What We Preserve
Our archive focuses on the foundational era of the World Wide Web (1990–1999), with selective preservation extending into the early 2000s. We prioritize historically significant domains, grassroots communities, experimental HTML, and culturally pivotal moments in digital history.
Verification & Fidelity Process
Every archived record undergoes a multi-stage verification pipeline designed to ensure pixel-accurate and structural fidelity to the original source at the time of capture.
Snapshot Capture
Full HTTP/1.0 & 1.1 transaction logging, including headers, cookies, and response streams. Stored in WARC 1.1/2.0 format.
Cryptographic Hashing
SHA-256 hashes generated for every payload. Cross-referenced with original server responses where available.
Structural Validation
HTML parsers verify DOCTYPE, frame sets, table layouts, and deprecated tags against era-specific W3C drafts.
Visual Regression
Headless Netscape/Naviator & IE4 renderers generate checksums of rendered output for spot-check verification.
Storage & Preservation Standards
We maintain strict bit-level preservation protocols to prevent digital decay and ensure long-term accessibility of fragile web artifacts.
| Standard | Implementation | Status |
|---|---|---|
| WARC Compliance | IETF RFC 8493 & RFC 8494 | Active |
| Checksum Verification | SHA-256 + MD5 fallback | Active |
| Redundancy | Geo-distributed triple mirroring | Active |
| Bit-Rot Detection | Quarterly scrubbing & parity repair | Active |
| Format Migration | Auto-conversion to next-gen archival specs | Pending Review |
How We Capture the Web
Our crawling strategy balances historical completeness with ethical preservation. We operate within documented historical boundaries while respecting early web infrastructure limitations.
Scope Rules: We prioritize domains registered between 1990–1999, GeoCities/Angelfire/Tripod networks, university `.edu` nodes, and government `.gov` portals. Commercial sites are archived based on cultural impact and technological novelty.
Rate Limiting: Historical crawling respects legacy server capacities. We use 1–3 second delays per request and queue-based throttling to prevent overwhelming legacy infrastructure during recovery attempts.
Dynamic Content: JavaScript-heavy pages are captured via emulated 1995–1999 browser environments (Netscape 3/4, IE4/5, Opera 3) to preserve interactivity where possible.
Audit Logs & Researcher Access
Accuracy requires transparency. We publish monthly integrity reports and maintain open audit trails for all archival operations.
- Capture Metadata: Timestamps, User-Agent strings, HTTP response codes, and redirect chains
- Hash Verification Logs: Publicly accessible cryptographic proof for every archived payload
- Gap Reporting: Documented lists of unreachable, expired, or legally restricted content
- Researcher API: Batch download endpoints with integrity validation flags for academic use