2026-09-05 06:07:25 +05:30
2026-09-05 06:07:25 +05:30
2026-09-05 06:07:25 +05:30
2026-09-05 06:07:25 +05:30
2026-09-05 06:07:25 +05:30

SiteHarbor

SiteHarbor is a standalone Python web-capture service. It creates a bounded offline archive of a public website, preserves discovered static assets, rewrites supported local links, records a manifest, and keeps completed ZIP files outside the public web root.

It is intentionally designed as a local or trusted-network application. It uses SQLite for durable job state and does not require Redis, a queue server, or object storage.

What It Does

Browser -> FastAPI -> SQLite job queue -> local capture worker
        -> private local ZIP and JSON report -> browser-scoped download route

The browser creates an opaque local session key on first use. It is stored in local storage and scopes capture status, cancellation, reports, and downloads to that browser. The worker and SQLite job continue independently when the page disconnects; reopening the page reconnects to locally recorded sessions. Clearing browser storage intentionally removes access to those sessions.

The capture engine:

  • Crawls same-site documents within configured page, depth, size, and time budgets.
  • Captures linked CSS, JavaScript, image, font, media, manifest, SVG, iframe, and CSS-import resources.
  • Can include linked third-party assets while avoiding third-party document crawling.
  • Rewrites supported HTML, CSS, and static JavaScript references to local archive paths.
  • Produces siteharbor-manifest.json and README.txt inside each archive.
  • Reports skipped resources, limits, failures, redirects, and static-capture caveats.

No static crawler can guarantee a fully functioning offline copy of every modern web application. Dynamic API calls, authenticated content, service workers, CAPTCHAs, backend state, and runtime-generated resources need a separately isolated browser-rendered mode in a later release.

Quick Start

Use Python 3.11 or newer.

py -3.11 -m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install --upgrade pip
python -m pip install -e .[dev]
siteharbor

Open http://127.0.0.1:8787.

To use Uvicorn directly:

python -m uvicorn app.main:app --host 127.0.0.1 --port 8787

The default bind address is loopback. Do not expose SiteHarbor directly to the public Internet without authentication, a network egress policy, rate limits, and operational controls.

Configuration

All configuration uses environment variables prefixed with SITEHARBOR_.

Variable Default Purpose
SITEHARBOR_HOST 127.0.0.1 Bind address used by siteharbor.
SITEHARBOR_PORT 8787 Bind port used by siteharbor.
SITEHARBOR_DATA_DIR ./data SQLite database, work directories, artifacts, and reports.
SITEHARBOR_WORKER_CONCURRENCY 1 Capture workers per application process.
SITEHARBOR_FETCH_CONCURRENCY 6 Maximum parallel fetches a visitor may select for one capture.
SITEHARBOR_LEASE_SECONDS 120 SQLite worker lease duration; values below 15 seconds are rejected.
SITEHARBOR_ALLOW_PRIVATE_NETWORKS false Development-only override that permits loopback and private hosts.
SITEHARBOR_ALLOW_NONSTANDARD_PORTS false Development-only override for ports other than 80 and 443.
SITEHARBOR_RESPECT_ROBOTS true Respect robots rules during capture.
SITEHARBOR_DEFAULT_RETENTION_DAYS 7 Default retention period for completed artifacts.
SITEHARBOR_PROXY_URL unset Operator-approved HTTP(S) proxy route. Visitors can only opt into this configured route; they cannot supply proxy endpoints.

For a local fixture site on localhost:8000, use both development overrides:

$env:SITEHARBOR_ALLOW_PRIVATE_NETWORKS = "true"
$env:SITEHARBOR_ALLOW_NONSTANDARD_PORTS = "true"
siteharbor

SQLite and Uvicorn Workers

SiteHarbor uses SQLite WAL mode and a lease-based database claim to run jobs without Redis. A single Uvicorn process with SITEHARBOR_WORKER_CONCURRENCY=1 is the recommended standalone configuration.

Multiple Uvicorn processes can claim jobs from the same local SQLite database, but SQLite still permits only one writer at a time. Use modest concurrency and shared local storage only. This design is intentionally for standalone or small trusted deployments, not high-scale public crawling.

Safety Boundaries

The application blocks non-HTTP(S) targets, URL credentials, nonstandard ports by default, and private/reserved network destinations. It revalidates and pins each request target to validated numeric addresses during crawling and redirects, avoiding an uncontrolled second DNS lookup at connection time.

If the service is ever exposed beyond localhost, still route worker traffic through a controlled public-only egress gateway and add authentication, quotas, and rate limiting. The standalone safeguards are intentionally strong, but a network boundary remains defense in depth for a public service.

Browser Sessions and Public Metrics

  • No account, login, or signup flow is required.
  • The public page exposes aggregate counts only: successful captures, data collected, active crawls, and files preserved.
  • Private capture routes require the opaque browser-local session key. A different browser cannot list, inspect, cancel, delete, download, or open events for another browser's captures.
  • This local session key provides browser ownership, not public-service abuse protection. A public deployment still needs rate limits, quotas, egress controls, and monitoring.

Development Commands

python -m pytest
python -m ruff check .
python -m compileall app
S
Description
Build an offline-ready snapshot of a public site. The service runs the crawl; this browser keeps the key to its own sessions.
Readme
177 KiB
Languages
Python 70.6%
JavaScript 13.1%
CSS 10.5%
HTML 5.8%