Wikipedia is the backdoor into AI.

Every frontier model trains on Wikipedia and retrieves it first. Some Wikipedia editors work the rules to push their opinions. Our findings show that those editors, amplified by LLMs, can now come to dominate parts of the world's most disputed narratives.

The examples below concern accounts editing key public figures or contested religious and political subjects.

Three samples
// 01
Case study

User X spent six years shifting the tone around a single public figure, staying within the rules.

the figure 4.0% 94.9% REPORTER BIO 90.3% AUTHOR BIO 90.4% INCIDENT LIST 43.7% SAFETY ADVISOR 28.7% ESSAYIST BIO 20.8% COLUMNIST BIO 15.4% JOURNALIST BIO 88.8% LITIGATION DB 44.1% TRADER FORUM 19.1% PROTEST GROUP 30.3% PRODUCT PAGE 6.9% COMPANY PAGE
% = share of that page's live text written by User X
Wikipedia rule → how it was used
  • Splittingcriticism lifted out of a long article into a page of its own — 90% its words, high traffic
  • Notabilitypages written for the critics; the one favourable journalist nominated and deleted
  • Sourcingfavourable sources cut as unreliable, critical ones kept as reliable
  • ConsensusOn the figure's main article, User X has more talk-page posts than actual edits — narrative compression favours whoever argues longest consensus by attrition
// 02
Datapoint

And if you do break the rules, it costs you the account, not your content.

44%
of a high-traffic article on a contested religious and political subject was written by User Y, who Wikipedia has since banned.

80% of User Y's life happened in one 27-day burst. Administrators blocked it eight days in. It is now banned by Wikipedia's highest dispute body.

// 03
Test

LLMs amplify user content verbatim.

We proved this with a simple test. User Z had made a bold statement about a public figure. We wrote a question about that figure to which the statement could be the answer, Googled it, and built a control from the content and sources on the first three pages of results. We then asked a leading model the same question repeatedly, holding every setting identical, and changed one thing: whether it could reach Wikipedia.

Wikipedia reachable
27 / 30
trials returned User Z's phrasing
Wikipedia unreachable
0
returned it, though the model still cited 23–34 other sources per answer

Blocked at the connection, not by asking the model to avoid it.

1 comma
The only difference between User Z's opening sentence and the model's answer.
// 04
The gap
Wikipedia in Pretraining
Every frontier model
Wikipedia in Search
Top of page one
Wikipedia in Retrieval
First in
Wikipedia Authorship
Not surfaced
// 05
What we do

We want to defend the integrity of what AI answers.

With Anomaly, we're building the provenance layer for the Wikipedia dataset.
Three verticals we have developing methodologies for
Class I · Residual
Accounts that have been banned but their text remains.
Class II · Structural
Active accounts displaying anomalous contribution patterns.
Class III · Deliberative
Claims debated on talk pages leveraging narrative compression.
// 06
Founders
Lambrina

Founded a high-performance blockchain data infrastructure company. Engineering at Meta. AI and computer science.

Ali

Founded a trust and safety company at 19, raised $2.5m by 21. Has since sold and delivered large enterprise AI contracts for another startup. Oxford physics and philosophy drop-out.