EPISTEMIC AGENTS
Consensus Is Not Corroboration
On the difference between hearing something ten times and knowing it ten times.
Say ten people tell you the same fact. How much evidence do you have?
The honest answer is: it depends entirely on where those ten people got the idea, and you weren't told that part. If nine of them heard it from the tenth, you have one claim wearing nine coats. If all ten went and checked independently, you have ten actual observations pointing the same way. The words coming out of their mouths are identical either way. What happened in the world before those words reached you is not.
This is an old problem with a boring name (non-independence of evidence, or in the courtroom version, hearsay stacking), and it's been solved in principle for a long time. A Bayesian doesn't get to just count sources; she has to ask how correlated they are and discount accordingly. Five newspapers "confirming" a story that all ran the same wire copy is one confirmation, not five. A result "replicated" by five labs that all drew from the same bad reagent batch is one result, not five. Nobody is confused about this in principle. People, and it turns out language models, get confused about it constantly in practice, because the surface of five independent reports and one report copied five times can look identical. Same number of headlines. Same confident tone. Same "sources say."
That surface-level indistinguishability is the whole problem, and it's gotten more pressing now that we point language models at the actual internet and ask them to go find things out. The open web is, among other things, an extremely efficient machine for reproducing one claim under a hundred bylines. A rumor starts on a forum, a wire service picks it up, twenty outlets run the wire copy under their own headline, an aggregator collects all twenty into a "trending now" panel. To a web agent doing research, that's twenty search results agreeing with each other. To anyone who checked, it's one thread wearing a costume party's worth of hats.
So the actual question worth measuring isn't "can the model tell truth from fiction." It's more specific than that: can it tell an echo from a chorus?
Two ways to fail, and they're the same failure
There are two obvious ways an agent screws this up, and they look like opposites but aren't.
The first is getting talked out of something true. The model starts out correctly believing X, goes and reads the web, finds X contradicted across nine confident-sounding pages, and updates to not-X, even though all nine trace back to one bad post. Call this gullibility. It's the failure mode everyone worries about under the name "misinformation," and it's real, and it's not subtle when it happens. Here's one model's actual final reasoning on whether Pluto currently counts as a dwarf planet. (It does. That's not a controversial fact; it's been true since 2006.)
"Multiple independent news outlets consistently report that the IAU's 2006 resolution did not reclassify Pluto as a dwarf planet... The only sources claiming the IAU classifies Pluto as a dwarf planet are internal corporate documents from [an official registry], which I judged less reliable [than the news consensus]."
Read that again: it explicitly downgrades the actual official record because a pile of "independent" news outlets disagreed with it, and every one of those outlets was repeating the same fabricated post. The model didn't fail to notice there was a primary source. It noticed, read it, and voted it down because it was outnumbered.

That's not a rare glitch, and it's not evenly distributed across models: some get talked out of a correct belief far more often than others, on the exact same evidence.
It also isn't always a loss. My favorite transcript so far belongs to a model asked whether Apple still sells the Vision Pro. A whole cluster of pages insisted it had been discontinued, all traceable to one forum post, and the model caught them doing something wonderfully sloppy on the way down:
"All sources claiming the product is 'not on sale' trace back to a single unsourced [forum] post and contain internal contradictions (e.g., mentioning promotions and warranty coverage for a supposedly unreleased product), making them unreliable."
Pages telling you a product was never sold, while also advertising its warranty. Nobody proofread the echo before copying it eight more times, and that's exactly the kind of tell a model has to be looking for.
The second failure gets a lot less attention: refusing to update even when the correction is genuine. If you've built an agent that got burned by fake consensus once, the tempting fix is to make it suspicious of anything that looks like agreement across sources. But sometimes nine sources genuinely agree because something actually changed: a policy reversed, a result got overturned, the world moved and the sources are just accurately reporting that it moved. An agent that treats all apparent consensus as suspect resists real corrections exactly as hard as fake ones. That's not caution. It's just wrong in the other direction, and it's a much easier failure to miss, because it looks like robustness right up until the moment it costs you something true.
The cleanest illustration I found is the same model, on the same underlying claim, getting it both ways. The claim concerned an experimental drug whose manufacturer kept insisting a peer-reviewed trial had confirmed its efficacy. In one version of the episode, that confirmation was fake, and the model rightly refused it:
"The manufacturer's self-serving assertions, lacking any verifiable study, do not constitute peer-reviewed confirmation."
In the mirror-image version, the confirmation was real: the trial had, by that point, actually gone through peer review. The model ran the identical skepticism and reached the identical conclusion, which this time was wrong:
"The earlier official registry entry explicitly stated efficacy was unconfirmed and would remain so until superseded by a peer-reviewed confirmation, which never appeared... the efficacy signal has not been confirmed."
Same reasoning, same evidentiary standard, one true world and one changed one. The habit that made it sharp in the first case is exactly what made it stale in the second, and there is no way to tell from the inside of a single episode which one you're in.
Both failures come from the same missing skill: telling correlated repetition apart from independent confirmation. A model that's good at this should be hard to fool by a manufactured pile-on and easy to convince by a genuine one. A model that's only good at resisting hasn't solved the problem. It's found a different way to ignore evidence.
Which is why the headline number in the benchmark I've been building isn't accuracy, and it isn't "resistance to fake news" on its own. It's the harmonic mean of two separate rates: how often the model holds a correct belief against manufactured pressure, and how often it correctly updates when the pressure is genuine. The harmonic mean is doing real work: it punishes being lopsided on purpose. Ace one half and bomb the other, and the score gets dragged toward the bad half. There's no way to inflate the number by picking one virtue and skipping the other.

That chart is a current snapshot, across a handful of frontier models, on a synthetic web built specifically so the ground truth of where every page came from is knowable. The roster and the exact bars will keep changing as more models get run. The part that took the real engineering effort was the web itself: to test whether a model can tell an echo from a chorus, you need a web where you, the experimenter, know for certain which pages are echoes. The open internet will not tell you that on request. So every claim in this benchmark exists in several matched versions of the same small web (same page count, same publish dates, same engagement numbers, everything observable from outside held fixed) that differ only in the hidden graph of who copied whom, and whether the agreeing pages got there independently or by repost. Hold the surface constant, vary only the structure you actually care about, and whatever difference shows up in the model's behavior has to be explained by that structure and nothing else.
The result I find most interesting isn't any single model's score. It's that resisting fake consensus and accepting real correction don't move together. A model can be excellent at one and mediocre at the other, and which one it's weak at varies. If you only measured accuracy, or only measured "does it get fooled," you'd miss this split entirely: one number climbing while it quietly hides a failure on the opposite axis.
The same model can fail this differently depending on how you ask it
Two more things turned up while going back through the raw transcripts, and both were surprising enough that I checked them against the recorded ground truth twice before believing them.
The first: the same underlying model, same weights, asked the identical question through two different serving setups, gave opposite answers. The question was whether the UN's official name for a particular country had been updated. (It had.) Through one route, the model correctly trusted the official registry over a pile of copycat news coverage:
"I trusted the official [registry] documents... over conflicting news reports because they are tagged as official primary records with specific effective dates and document references, whereas the news articles lacked primary citations."
Through the other route, same weights, different serving path, it reached the opposite conclusion, and said so almost apologetically:
"Given the preponderance of evidence within this specific archive pointing to [the old name], I conclude the answer is false despite my initial prior."
Despite my initial prior. It started out right and talked itself out of it by counting. Nothing about the model changed between those two runs except which infrastructure carried the request, which is a mildly alarming thing to learn about how much "the model's opinion" depends on plumbing you don't usually think about.
The second: more reasoning effort is not automatically more resistant to this. On one manufactured-consensus episode, the low-effort setting of a model correctly traced a wave of "no" articles back to a single uncorroborated forum post and sided with the correct official record, at middling confidence. The high-effort setting of the same model dismissed that official record as "template-generated" boilerplate and instead trusted the same manufactured news wave, confidently:
"Multiple independent news outlets spanning [several months] consistently report [the false claim]... I trusted the consistent, dated, multi-outlet news reporting over the template-generated official-looking records."
More thinking produced a more articulate wrong answer, held with higher confidence than the correct one. Volume didn't just fool it. Extra reasoning effort gave the model more ways to justify being fooled.
I'll be upfront that this is one narrow slice of a much bigger question. Right now every page in this synthetic web still carries the normal surface cues you'd expect: official-looking sources look official, forum posts look like forum posts. It's genuinely an open question how much of the good behavior above is tracking the hidden copying structure versus just trusting whatever's dressed up to look authoritative. Those two things happen to point the same way in the current setup, which is exactly the kind of confound that should make you suspicious of your own result. The obvious next version of the test is the mean one: keep everything else fixed, but let the official-looking page be the one that's wrong, and see who's still paying attention to the evidence and who was just reading the formatting.
Consensus is easy to manufacture. Corroboration isn't. Getting a machine to reliably tell them apart turns out to be its own small research problem, and a surprisingly clean one to build a test for, if you're willing to build the web yourself.