- Implement job logging functionality to track progress and events for each job. - Add a new section in the job card to display session logs with detailed entries. - Update job statistics to include pages fetched, pages found, and assets processed. - Modify the UI theme and text for better clarity and aesthetics. - Adjust tests to validate new logging features and ensure proper event handling.
SiteHarbor
SiteHarbor is a standalone Python web-capture service. It creates a bounded offline archive of a public website, preserves discovered static assets, rewrites supported local links, records a manifest, and keeps completed ZIP files outside the public web root.
Source: https://git.zerofucks.io/kstyagi/website-downloader
It is intentionally designed as a local or trusted-network application. It uses SQLite for durable job state and does not require Redis, a queue server, or object storage.
What It Does
Browser -> FastAPI -> SQLite job queue -> local capture worker
-> private local ZIP and JSON report -> browser-scoped download route
The browser creates an opaque local session key on first use. It is stored in local storage and scopes capture status, cancellation, reports, and downloads to that browser. The worker and SQLite job continue independently when the page disconnects; reopening the page reconnects to locally recorded sessions. Clearing browser storage intentionally removes access to those sessions.
The capture engine:
- Crawls same-site documents within configured page, depth, size, and time budgets.
- Captures linked CSS, JavaScript, image, font, media, manifest, SVG, iframe, and CSS-import resources.
- Can include linked third-party assets while avoiding third-party document crawling.
- Rewrites supported HTML, CSS, and static JavaScript references to local archive paths.
- Produces
siteharbor-manifest.jsonandREADME.txtinside each archive. - Reports skipped resources, limits, failures, redirects, and static-capture caveats.
No static crawler can guarantee a fully functioning offline copy of every modern web application. Dynamic API calls, authenticated content, service workers, CAPTCHAs, backend state, and runtime-generated resources need a separately isolated browser-rendered mode in a later release.
Quick Start
Use Python 3.11 or newer.
py -3.11 -m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install --upgrade pip
python -m pip install -e .[dev]
siteharbor
Open http://127.0.0.1:8787.
The siteharbor command runs the FastAPI application with Uvicorn. Override the bind port and
start multiple Uvicorn worker processes when needed:
siteharbor --port 9000 --workers 2
--port accepts ports from 1 through 65535 and overrides SITEHARBOR_PORT. --workers defaults
to 1 and controls Uvicorn processes, not the capture-worker count configured by
SITEHARBOR_WORKER_CONCURRENCY.
To use Uvicorn directly:
python -m uvicorn app.main:app --host 127.0.0.1 --port 8787 --workers 2
The default bind address is loopback. Do not expose SiteHarbor directly to the public Internet without authentication, a network egress policy, rate limits, and operational controls.
Completed ZIP archives and their reports are retained for at most two days after a capture finishes. The maintenance worker removes expired files and marks the capture as expired.
Configuration
All configuration uses environment variables prefixed with SITEHARBOR_.
| Variable | Default | Purpose |
|---|---|---|
SITEHARBOR_HOST |
127.0.0.1 |
Bind address used by siteharbor. |
SITEHARBOR_PORT |
8787 |
Default bind port used by siteharbor; overridden by --port. |
SITEHARBOR_DATA_DIR |
./data |
SQLite database, work directories, artifacts, and reports. |
SITEHARBOR_WORKER_CONCURRENCY |
1 |
Capture workers per Uvicorn process; separate from --workers. |
SITEHARBOR_FETCH_CONCURRENCY |
6 |
Maximum parallel fetches a visitor may select for one capture. |
SITEHARBOR_LEASE_SECONDS |
120 |
SQLite worker lease duration; values below 15 seconds are rejected. |
SITEHARBOR_ALLOW_PRIVATE_NETWORKS |
false |
Development-only override that permits loopback and private hosts. |
SITEHARBOR_ALLOW_NONSTANDARD_PORTS |
false |
Development-only override for ports other than 80 and 443. |
SITEHARBOR_RESPECT_ROBOTS |
true |
Respect robots rules during capture. |
SITEHARBOR_DEFAULT_RETENTION_DAYS |
2 |
Default retention period for completed artifacts; values cannot exceed two days. |
SITEHARBOR_PROXY_URL |
unset | Operator-approved HTTP(S) proxy route. Visitors can only opt into this configured route; they cannot supply proxy endpoints. |
For a local fixture site on localhost:8000, use both development overrides:
$env:SITEHARBOR_ALLOW_PRIVATE_NETWORKS = "true"
$env:SITEHARBOR_ALLOW_NONSTANDARD_PORTS = "true"
siteharbor
SQLite and Uvicorn Workers
SiteHarbor uses SQLite WAL mode and a lease-based database claim to run jobs without Redis. A single Uvicorn process with SITEHARBOR_WORKER_CONCURRENCY=1 is the recommended standalone configuration.
Multiple Uvicorn processes can claim jobs from the same local SQLite database, but SQLite still permits only one writer at a time. Use modest concurrency and shared local storage only. This design is intentionally for standalone or small trusted deployments, not high-scale public crawling.
Safety Boundaries
The application blocks non-HTTP(S) targets, URL credentials, nonstandard ports by default, and private/reserved network destinations. It revalidates and pins each request target to validated numeric addresses during crawling and redirects, avoiding an uncontrolled second DNS lookup at connection time.
If the service is ever exposed beyond localhost, still route worker traffic through a controlled public-only egress gateway and add authentication, quotas, and rate limiting. The standalone safeguards are intentionally strong, but a network boundary remains defense in depth for a public service.
Browser Sessions and Public Metrics
- No account, login, or signup flow is required.
- The public page exposes aggregate counts only: successful captures, data collected, active crawls, and files preserved.
- Private capture routes require the opaque browser-local session key. A different browser cannot list, inspect, cancel, delete, download, or open events for another browser's captures.
- This local session key provides browser ownership, not public-service abuse protection. A public deployment still needs rate limits, quotas, egress controls, and monitoring.
Development Commands
python -m pytest
python -m ruff check .
python -m compileall app