How the archives are built.
Every vertical follows the same five steps: collect, gate, extract, preserve and index. The details change from subject to subject. The shape doesn't.
Collecting the sources
- In
- The source list
- Out
- Collected items
Each vertical has a fixed list of sources that it declares openly. The list is written down before collecting starts, and every change to it is recorded.
- News agencies (wire services)
- National broadcasters
- State-run media
- Feeds from institutions
- Specialist newsletters
- Social media accounts
Picked so the same event is covered by people who disagree about it.
- Prediction markets and price data
- Hourly
- News agency and broadcaster feeds
- Every few hours
- Reports from institutions
- Daily
- Media on messaging platforms Downloaded straight away. The download links expire within hours, and a web address saved without its file is a record of nothing.
- As collected
Sources are checked on a schedule, not swept up whenever something turns up. How often depends on how fast the subject moves and how fast the source rewrites itself.
Anyone can see what was looked at, and what wasn't.
If coverage changes over time, it's because the world changed, not because the collector did.
Keeping what's relevant
- In
- Collected items
- Out
- Relevant items
- Aside
- Rejects, logged
Most of what a source publishes has nothing to do with the subject. So every item goes through a weighted filter, the gate, before it's stored.
-
Matches the subject
It's checked for keywords, and for the names of people, places and organisations.
-
Recent enough
It has to fall inside a time window.
-
Long enough
It has to reach a minimum length. Anything shorter is a stub, not a document.
Each one is kept with its reason. That record is the only way to answer the question that matters about a filter: not what did it keep, but what did it throw away, and was that a mistake.
Sources judged to be first-hand voices on the subject skip the topic filter and are archived in full, because for them the way they frame things is the data.
Pulling out the text
- In
- Relevant items
- Out
- Text and details
- Aside
- Failures, kept
Each stored item is boiled down to its full text plus a set of details about it.
-
It doesn't work every time
Paywalls, pages that only build themselves in a browser, and deliberately awkward page code all cause failures. The success rate is tracked and published, not quietly hidden.
-
A failure is kept as a failure
A failed attempt isn't quietly deleted. A gap you can see can be fixed later. A gap you can't see can't.
Keeping a copy
- In
- Text and details
- Out
- Two copies
Collecting early matters because what was collected stops matching what was published. Articles get retitled, softened, corrected without notice, and removed.
The archive holds what was collected, exactly as it was when it was collected. If a source changes a document later, the archive doesn't rewrite its copy to match. The difference between the two is often the most revealing thing in the record.
-
A copy held by someone else
Where possible, items are also sent to the Internet Archive at the moment they're collected. That's a third-party copy, and it doesn't depend on this project still existing.
Sorting and showing
- In
- Saved copies
- Out
- A site to browse
Documents are sorted by topic and added to an index so they can be found. Websites are built on top of that index.
-
Sorted by a local AI model
A locally run AI model tags each document by topic, using an evolving set of categories defined for each vertical.
-
What the sites show
A collection you can browse, an overview of how much was collected and what it covers over time and, where a vertical has one, a search or summary tool.
A tag tells you where to look.
A tag can't tell you what's true. Tagging is done by machine, and it's wrong some of the time.
How up to date it is
- In
- Everything stored
- Out
- A dated website
Collecting runs all the time. Publishing doesn't.
-
The site lags behind
Each vertical rebuilds its website on a schedule, so what you see can be up to one publishing cycle behind what's been collected.
-
Every site is stamped
Each vertical carries a provenance stamp: a label giving the time its site was built and the figures at that moment.
-
Read the stamp literally
It shows the state of the collection at that time, frozen into a page that hasn't changed since. It isn't a live connection to anything. On the front page, the meter reading and the date beside each station come from the same stamp.