104 lines
5.5 KiB
Markdown
104 lines
5.5 KiB
Markdown
# SiteHarbor
|
|
|
|
SiteHarbor is a standalone Python web-capture service. It creates a bounded offline archive of a public website, preserves discovered static assets, rewrites supported local links, records a manifest, and keeps completed ZIP files outside the public web root.
|
|
|
|
It is intentionally designed as a local or trusted-network application. It uses SQLite for durable job state and does not require Redis, a queue server, or object storage.
|
|
|
|
## What It Does
|
|
|
|
```text
|
|
Browser -> FastAPI -> SQLite job queue -> local capture worker
|
|
-> private local ZIP and JSON report -> browser-scoped download route
|
|
```
|
|
|
|
The browser creates an opaque local session key on first use. It is stored in local storage and
|
|
scopes capture status, cancellation, reports, and downloads to that browser. The worker and SQLite
|
|
job continue independently when the page disconnects; reopening the page reconnects to locally
|
|
recorded sessions. Clearing browser storage intentionally removes access to those sessions.
|
|
|
|
The capture engine:
|
|
|
|
- Crawls same-site documents within configured page, depth, size, and time budgets.
|
|
- Captures linked CSS, JavaScript, image, font, media, manifest, SVG, iframe, and CSS-import resources.
|
|
- Can include linked third-party assets while avoiding third-party document crawling.
|
|
- Rewrites supported HTML, CSS, and static JavaScript references to local archive paths.
|
|
- Produces `siteharbor-manifest.json` and `README.txt` inside each archive.
|
|
- Reports skipped resources, limits, failures, redirects, and static-capture caveats.
|
|
|
|
No static crawler can guarantee a fully functioning offline copy of every modern web application. Dynamic API calls, authenticated content, service workers, CAPTCHAs, backend state, and runtime-generated resources need a separately isolated browser-rendered mode in a later release.
|
|
|
|
## Quick Start
|
|
|
|
Use Python 3.11 or newer.
|
|
|
|
```powershell
|
|
py -3.11 -m venv .venv
|
|
.\.venv\Scripts\Activate.ps1
|
|
python -m pip install --upgrade pip
|
|
python -m pip install -e .[dev]
|
|
siteharbor
|
|
```
|
|
|
|
Open `http://127.0.0.1:8787`.
|
|
|
|
To use Uvicorn directly:
|
|
|
|
```powershell
|
|
python -m uvicorn app.main:app --host 127.0.0.1 --port 8787
|
|
```
|
|
|
|
The default bind address is loopback. Do not expose SiteHarbor directly to the public Internet without authentication, a network egress policy, rate limits, and operational controls.
|
|
|
|
## Configuration
|
|
|
|
All configuration uses environment variables prefixed with `SITEHARBOR_`.
|
|
|
|
| Variable | Default | Purpose |
|
|
| --- | --- | --- |
|
|
| `SITEHARBOR_HOST` | `127.0.0.1` | Bind address used by `siteharbor`. |
|
|
| `SITEHARBOR_PORT` | `8787` | Bind port used by `siteharbor`. |
|
|
| `SITEHARBOR_DATA_DIR` | `./data` | SQLite database, work directories, artifacts, and reports. |
|
|
| `SITEHARBOR_WORKER_CONCURRENCY` | `1` | Capture workers per application process. |
|
|
| `SITEHARBOR_FETCH_CONCURRENCY` | `6` | Maximum parallel fetches a visitor may select for one capture. |
|
|
| `SITEHARBOR_LEASE_SECONDS` | `120` | SQLite worker lease duration; values below 15 seconds are rejected. |
|
|
| `SITEHARBOR_ALLOW_PRIVATE_NETWORKS` | `false` | Development-only override that permits loopback and private hosts. |
|
|
| `SITEHARBOR_ALLOW_NONSTANDARD_PORTS` | `false` | Development-only override for ports other than 80 and 443. |
|
|
| `SITEHARBOR_RESPECT_ROBOTS` | `true` | Respect robots rules during capture. |
|
|
| `SITEHARBOR_DEFAULT_RETENTION_DAYS` | `7` | Default retention period for completed artifacts. |
|
|
| `SITEHARBOR_PROXY_URL` | unset | Operator-approved HTTP(S) proxy route. Visitors can only opt into this configured route; they cannot supply proxy endpoints. |
|
|
|
|
For a local fixture site on `localhost:8000`, use both development overrides:
|
|
|
|
```powershell
|
|
$env:SITEHARBOR_ALLOW_PRIVATE_NETWORKS = "true"
|
|
$env:SITEHARBOR_ALLOW_NONSTANDARD_PORTS = "true"
|
|
siteharbor
|
|
```
|
|
|
|
## SQLite and Uvicorn Workers
|
|
|
|
SiteHarbor uses SQLite WAL mode and a lease-based database claim to run jobs without Redis. A single Uvicorn process with `SITEHARBOR_WORKER_CONCURRENCY=1` is the recommended standalone configuration.
|
|
|
|
Multiple Uvicorn processes can claim jobs from the same local SQLite database, but SQLite still permits only one writer at a time. Use modest concurrency and shared local storage only. This design is intentionally for standalone or small trusted deployments, not high-scale public crawling.
|
|
|
|
## Safety Boundaries
|
|
|
|
The application blocks non-HTTP(S) targets, URL credentials, nonstandard ports by default, and private/reserved network destinations. It revalidates and pins each request target to validated numeric addresses during crawling and redirects, avoiding an uncontrolled second DNS lookup at connection time.
|
|
|
|
If the service is ever exposed beyond localhost, still route worker traffic through a controlled public-only egress gateway and add authentication, quotas, and rate limiting. The standalone safeguards are intentionally strong, but a network boundary remains defense in depth for a public service.
|
|
|
|
## Browser Sessions and Public Metrics
|
|
|
|
- No account, login, or signup flow is required.
|
|
- The public page exposes aggregate counts only: successful captures, data collected, active crawls, and files preserved.
|
|
- Private capture routes require the opaque browser-local session key. A different browser cannot list, inspect, cancel, delete, download, or open events for another browser's captures.
|
|
- This local session key provides browser ownership, not public-service abuse protection. A public deployment still needs rate limits, quotas, egress controls, and monitoring.
|
|
|
|
## Development Commands
|
|
|
|
```powershell
|
|
python -m pytest
|
|
python -m ruff check .
|
|
python -m compileall app
|
|
```
|