A comprehensive strategy for digital archaeology, institutional collaboration, and long-term preservation of early internet heritage. We are building the definitive record of web history before it fades.
Guiding principles that shape every technical and operational decision.
To systematically capture, preserve, and provide authenticated access to web content created between 1990 and 1999, ensuring that the cultural, technical, and historical record of the internet's formative decade remains intact for researchers, educators, and the public.
A globally accessible, academically recognized digital archive where no significant piece of early web history is lost to link rot, server decommissioning, or format obsolescence. We envision an open ecosystem where preservation is standardized, transparent, and sustainable.
Five core domains driving our preservation methodology and operational roadmap.
Automated crawling paired with targeted collection campaigns focusing on GeoCities, Angelfire, early commercial sites, institutional portals, and niche community hubs.
Recreating period-accurate rendering environments using containerized browsers, font preservation, and asset restoration to maintain visual and functional authenticity.
Formal partnerships with universities, libraries, museums, and tech historians to validate collections, share metadata, and co-develop preservation standards.
Geographically distributed storage, cryptographic verification (W3C WACZ format), automated integrity checks, and format migration pipelines to prevent digital decay.
Developer APIs, academic datasets, and public portals enabling computational analysis, historical research, and educational curriculum development.
Transparent copyright navigation, opt-out mechanisms, contextual attribution, and adherence to FAIR data principles while balancing public interest and creator rights.
Phased execution model with measurable milestones and deliverables.
Establish OAIS-compliant storage architecture, deploy targeted crawlers for high-risk domains, and complete initial digitization of 2M+ pages from 1994–1996.
Expand to 50+ institutional partners, launch researcher portal, implement AI-assisted metadata extraction, and secure long-term funding through grants and endowments.
Achieve 95% coverage of target domains, launch interactive timeline explorer, integrate with IIIF and W3C standards, and transition to community-governed stewardship.
Key performance indicators tracking preservation efficacy and ecosystem growth.
Ensuring long-term viability through transparent oversight and diversified funding.
Whether you're an institution, researcher, developer, or early web enthusiast, there's a role for you in safeguarding digital history.