Preserving public AI models, so the future of intelligence can interpret its past.
Project Pergamon is an archival effort to preserve publicly-distributed Large Language Models (LLMs) as part of the historical record.
We are building a professional archival standard for public AI: a practical preservation effort with the know-how of a scholarly institution and the passion of a historic moment at risk of being lost.
The Need to Preserve Public AI
Pubicly-distributed AI models increasingly rival subscription-based distributions in capability, while offering greater freedom, economy, and privacy to users. But existing public repositories are fragile. They can be delisted, neglected, censored, fragmented across platforms, or rendered unusable when the surrounding technical context disappears.
Access to public AI models should not be gated by mercy of cloud platforms and corporate incentives. The collective memory of the future deserves more than black boxes and broken links. If public distribution represents a freer and more democratic relationship to technology, it deserves institutions that can steward it for the next generation of builders and scholars.
What we preserve
Project Pergamon exists to address the unique archival challenge posed by generative Artificial Intelligence models. We start from the proposition that LLMs are not traditional software: Where software's behavior can be interpreted from a single source that a human can read, no one, single document fully determines what a model will do. Then, we design our process around documenting the full set of materials that can help future researchers interpret a model's likely behaviors from durable archival records, including:
- model weights
- architecture and configuration
- tokenizer and related metadata
- provenance and integrity information
- documentation for future custodians
Pergamon's approach to archiving these model is designed to work with and alongside established digital preservation standards, including the OAIS Reference Model (ISO 14721), Stanford's LOCKSS distributed-preservation architecture, and the highest tier of the NDSA Levels of Digital Preservation. Our long-term goal is to establish a durable preservation infrastructure that improves the odds of future technical and historical interpretation, even in events of censorship, catastrophe, and institutional loss.
How we do it
Modern archival tape makes preservation of public frontier models more feasible than many people assume. Today, the Linear Tape-Open (LTO) format represents the enterprise-grade archival tape standard used in real preservation workflows. LTO is our first archival substrate because it is the proven, scaleable, energy-efficient solution for professional archives. Likewise, the well-supported and open-source Linear Tape File System (LTFS) makes it easy for Pergamon's archival contributions to interface with institutions around the world.
At this stage, the challenge of preservation is more institutional than technical: multiple independent copies, routine verification, and stewardship across jurisdictions. To clear the grounds for collaboration, we are designing a broadly compatable metadata format using the popular BagIt packaging standard (RFC 8493). The Pergamon format will enable archivists to capture the unique nature of generative AI while also remaining interoperable with the standards of existing software archives.
The Archimedes Series
Our first public archive proofs will appear under The Archimedes Series as an impressive but credible preservation milestone: a single LTO-9 archival tape that holds a significant sample of today's major publicly-distributed frontier models. The Archimedes tape will employ our custom ARCHI data structure (see technical documentation): an AI-native archival standard that enables custodians to reliably, honestly, and safelty provide the public with access to frontier intelligence.
To create a representative picture of this era in public AI, we will weigh individual candidates for inclusion along several dimensions:
- License and openness — from permissive (Apache 2.0, MIT) through copyleft to open-weight-but-restricted, so the legal history of the period survives alongside the technical one.
- Scale — from megabyte embedding models to trillion-parameter systems, because both ends are part of AI history record.
- Deployment footprint — the models that are widely adopted for real world infrastructures and businesses, even if they don't attract headlines.
- Lineage — base models and the notable fine-tunes and distillations descended from them, preserving the genealogy of the field for our descendants.
- Origin — from frontier labs to independent builders, national initiatives to academic groups, across national and ideological divides, the corpus should reflect who made this era and where.
Interested colleagues and volunteers will be able to participate in the Archimedes selection process by suggesting and voting on models for inclusion. We are also seeking a technical collaborator to help co-author a whitepaper on the design of the first preservation pass and its path to scale.
Who we are
Project Pergamon was founded by Zachary Sheldon, a linguistic and cultural anthropologist. His current archival research program examines the deep history of algorithmic computing, with a focus on early modern Arabic mathematics and esotericism. His background spans preservation research and institution-building through organizing, non-profit collaboration, and, most recently, start-up AI work.
Zachary Sheldon built Project Pergamon with help from Aquarius (peopleaquarius.com), where he leads research on how people use AI-native tools to move complex projects from idea to execution.
Call for Collaborators
We are currently looking for collaborators in five areas:
- archival and library preservation
- technical preservation engineering
- AI model reconstruction and documentation
- legal and governance counsel
- fundraising and institutional development
Short version
Project Pergamon is building a global academic archive for publicly-distributed AI so models can be rebuilt after censorship, catastrophe, or institutional loss. We begin with tape not because one medium lasts forever, but because serious preservation starts with practical systems, honest documentation, and a commitment to carry knowledge forward.