SpinSignal
Failure Modes

How these archives can be wrong.

Everything below can genuinely go wrong with these archives. Some of it is handled. Some isn't, and it's listed anyway, because knowing where a record is weak beats being told it's strong.

Sources disappear

3 ways

Once something is published, it starts to decay, and it does so quietly.

  1. Links die

    URLs stop working. Outlets reorganise, drop their archives or go offline. A citation to a dead link isn't a citation.

  2. Articles get quietly rewritten

    Headlines change, figures get revised, paragraphs vanish, and often there's no correction. Two people reading “the same” article at different times may have read different articles.

  3. Sources go dark

    Accounts get suspended, channels get geoblocked, and some outlets become unreachable from certain networks mid-collection. Coverage drops for reasons that have nothing to do with events.

What's done

Collect early and often, store the full text rather than links, download media at collection time, and send a copy to a third-party archive where possible.

What isn't

Nothing fully protects against a source that vanished before it was captured. Gaps exist, and they aren't always detectable from inside the archive.

Machine errors

4 ways

The machinery makes its own mistakes. They're harder to spot than source errors, because the output looks clean either way.

  1. Made-up detail

    Where a vertical offers a generated summary, the model can state things its documents don't support, including fluent, specific, invented detail. Generated text is always the least trustworthy layer on any of these sites.

  2. Wrong attribution

    Syndication and nested quotes make it easy to file a claim under the wrong outlet, or to credit someone's words to the outlet quoting them. State media reprinting a wire story is a different signal from the wire service reporting it.

  3. Context gets cut off

    Retrieval works on passages, not whole documents. Lose the qualifier (a denial, an attribution, a “reportedly”) and a passage can mean the opposite while staying a word-for-word quote.

  4. Old data that looks current

    If collection stalls but publishing carries on, a site can render confidently over an index that stopped updating. That's close to invisible from outside, which is why every vertical stamps its generation time.

Model bias

5 ways

Language models do the extraction, tagging and summarising, and they aren't neutral.

  1. One model's blind spots

    One model's habits run through the whole corpus, so its errors are systematic, not random. Systematic errors don't average out with volume. They pile up.

  2. The wrong model for the subject

    Safety training can make some models evasive or unwilling on conflict casualties, atrocity claims and outbreak severity. A model that won't engage distorts the archive by leaving things out, and the gap is silent.

  3. Mostly English

    These pipelines run in English. Non-English sources are under-represented, and translation adds another lossy machine step. For conflicts fought largely in Persian, Hebrew and Arabic, that's a serious limitation, not a footnote.

  4. Training cutoff

    A model asked about events after its cutoff leans on what it knew before them. It can treat the unprecedented as normal and misjudge what matters.

  5. Choosing the sources

    The earliest bias is the hardest to undo: which sources go in the pool. It's a human editorial call, made once at the start, and everything downstream depends on it. Each vertical publishes its pool so you can inspect the choice.

Going stale

3 states

Every vertical declares a mode, and each one means something specific.

Live tracker
Collection runs on a schedule. The site is rebuilt from the store periodically, so it can trail the store by up to one publishing cycle. It never queries anything while you read.
Point-in-time archive
Collection has ended. The site is a snapshot of what was held on the stated date, and it won't grow. That's a legitimate end state, not a broken tracker.
Building
Announced, but not collecting yet. There's nothing to read and nothing to follow. On the front page these have no link, rather than a link to an empty site.
The dangerous case

A live tracker whose collection has quietly failed looks exactly like one that's working. Check the provenance stamp before trusting any figure as current. If the stamp is old, so is the page, however recently it was served to you.

Manipulation

4 ways

Some sources are in the pool because they're trying to shape the story. Their framing is evidence, so they stay. That has consequences.

  1. Disinformation is kept at equal weight

    The pipeline doesn't judge. A made-up casualty figure is stored with the same weight as a verified one, and nothing in the interface tells you which is which.

  2. Volume is not corroboration

    Coordinated networks repeat a claim across many accounts and outlets. An archive that counts occurrences will faithfully report that a claim is everywhere, which says nothing about whether it's true.

  3. A trusted source goes bad

    A source that's reliable when added may not stay that way. Compromise, capture or a change of ownership can change a feed's character without changing its URL.

  4. Indirect prompt injection

    Collected documents are fed to language models for tagging and summarising. Text written to be read as instructions rather than content is a live attack on any retrieval pipeline, this one included.

Misuse

Please read

This is independent research by one person and a set of automated pipelines. It isn't an intelligence product, journalism or advice, and it's not a basis for operational decisions.

  1. Don't use these archives to make claims about specific individuals.

  2. Don't treat aggregate counts as measurements of the world. They measure this collection.

  3. Don't cite a generated summary. Follow it to the underlying documents and cite those, once you've read them.

  4. If a vertical touches public health or an active conflict, the authorities with real reporting obligations are the right source for anything that matters.

Why this page exists

Why list all this? It's simple.

An archive that only advertises its strengths is asking to be trusted further than it can carry.

Knowing exactly where this one breaks is what makes the rest of it usable.