SpinSignal
Method

How the archives are built.

Every vertical follows the same five steps: collect, gate, extract, preserve and index. The details change from subject to subject. The shape doesn't.

Signal chain 5 stages · 3 outcomes · 2 copies
Many sources feed in at the start. Items the gate rejects, and items where pulling out the text fails, are set aside on lists instead of being lost: each one is logged with a reason, kept and counted. At the preserve step every item is copied out twice, to the Internet Archive and to the project's own store, and then carries on to the index.

Collecting the sources

In
The source list
Out
Collected items

Each vertical has a fixed list of sources that it declares openly. The list is written down before collecting starts, and every change to it is recorded.

What's on the list 6 kinds
  • News agencies (wire services)
  • National broadcasters
  • State-run media
  • Feeds from institutions
  • Specialist newsletters
  • Social media accounts

Picked so the same event is covered by people who disagree about it.

The schedule Set times
Prediction markets and price data
Hourly
News agency and broadcaster feeds
Every few hours
Reports from institutions
Daily
Media on messaging platforms Downloaded straight away. The download links expire within hours, and a web address saved without its file is a record of nothing.
As collected

Sources are checked on a schedule, not swept up whenever something turns up. How often depends on how fast the subject moves and how fast the source rewrites itself.

Two things matter more than the size of the list
Declared

Anyone can see what was looked at, and what wasn't.

Stable

If coverage changes over time, it's because the world changed, not because the collector did.

Keeping what's relevant

In
Collected items
Out
Relevant items
Aside
Rejects, logged

Most of what a source publishes has nothing to do with the subject. So every item goes through a weighted filter, the gate, before it's stored.

  1. Matches the subject

    It's checked for keywords, and for the names of people, places and organisations.

  2. Recent enough

    It has to fall inside a time window.

  3. Long enough

    It has to reach a minimum length. Anything shorter is a stub, not a document.

Every rejection is logged

Each one is kept with its reason. That record is the only way to answer the question that matters about a filter: not what did it keep, but what did it throw away, and was that a mistake.

First-hand voices skip the filter

Sources judged to be first-hand voices on the subject skip the topic filter and are archived in full, because for them the way they frame things is the data.

Pulling out the text

In
Relevant items
Out
Text and details
Aside
Failures, kept

Each stored item is boiled down to its full text plus a set of details about it.

  1. It doesn't work every time

    Paywalls, pages that only build themselves in a browser, and deliberately awkward page code all cause failures. The success rate is tracked and published, not quietly hidden.

  2. A failure is kept as a failure

    A failed attempt isn't quietly deleted. A gap you can see can be fixed later. A gap you can't see can't.

Keeping a copy

In
Text and details
Out
Two copies

Collecting early matters because what was collected stops matching what was published. Articles get retitled, softened, corrected without notice, and removed.

The original is what counts

The archive holds what was collected, exactly as it was when it was collected. If a source changes a document later, the archive doesn't rewrite its copy to match. The difference between the two is often the most revealing thing in the record.

  1. A copy held by someone else

    Where possible, items are also sent to the Internet Archive at the moment they're collected. That's a third-party copy, and it doesn't depend on this project still existing.

Sorting and showing

In
Saved copies
Out
A site to browse

Documents are sorted by topic and added to an index so they can be found. Websites are built on top of that index.

  1. Sorted by a local AI model

    A locally run AI model tags each document by topic, using an evolving set of categories defined for each vertical.

  2. What the sites show

    A collection you can browse, an overview of how much was collected and what it covers over time and, where a vertical has one, a search or summary tool.

Tags are for finding things

A tag tells you where to look.

Not for deciding what's true

A tag can't tell you what's true. Tagging is done by machine, and it's wrong some of the time.

How up to date it is

In
Everything stored
Out
A dated website

Collecting runs all the time. Publishing doesn't.

  1. The site lags behind

    Each vertical rebuilds its website on a schedule, so what you see can be up to one publishing cycle behind what's been collected.

  2. Every site is stamped

    Each vertical carries a provenance stamp: a label giving the time its site was built and the figures at that moment.

  3. Read the stamp literally

    It shows the state of the collection at that time, frozen into a page that hasn't changed since. It isn't a live connection to anything. On the front page, the meter reading and the date beside each station come from the same stamp.