A new paper out of Cambridge does something the multi-agent field has mostly avoided: it puts an exact number on what a debate between two AI models is worth. And the number, when the two models are cut from the same cloth, is zero.

Not "small." Zero. Robin Young's "Knowledge Divergence and the Value of Debate for Scalable Oversight" (arXiv:2603.05293, March 2026) proves that when two models share the same training data, making them argue produces exactly the same answer you would get from one of them thinking alone. The debate doesn't add a little. It adds nothing.

The disappointment everyone has felt

Anyone who has wired up a multi-agent debate knows the letdown. You set two models against each other, pay roughly twice the compute, watch them exchange paragraphs — and the final answer looks like what a single well-prompted model gives you. The field has names for why it sometimes goes wrong: sycophancy, conformity, models folding to whoever spoke last. What it has lacked is a theory of when debate should help at all.

Young supplies the missing condition, and it is not about how clever the models are. It is about how differently they know things.

There is empirical cover for this already. Goel and colleagues showed in 2025, in "Great models think alike and this undermines AI oversight," that as frontier models get more capable their mistakes become more correlated, and that diversity between models improves oversight. That was a measurement without a mechanism. Young's paper is the mechanism.

The reframe: debate runs on private information

The move is to stop treating the two debaters as abstract "agents" and start measuring how differently they represent the world. Young does this with principal angles between the models' representation subspaces — a geometric way of asking how much of what one model knows lives in a direction the other model simply cannot point at.

From there the whole value of debate collapses into a single quantity he calls the private information value: how much answer-relevant knowledge one model holds that the other's representation can't reach. When that quantity is zero, the debate advantage is provably zero. When it is large, debate isn't just helpful — it is the only way to get the answer.

Two copies of the same model don't debate. They agree with extra steps.

The scaling has a genuine phase transition. For models that are close to identical, the advantage is quadratically small — negligible, not worth the second inference bill. For models that are genuinely complementary, it grows large. The turn between those two worlds is sharp, not gradual.

Three regimes, and a pointed correction

Young sorts the possibilities into three cases. In the shared regime the models know the same things and debate buys nothing. In the one-sided regime one model already holds the better answer, and the debate structure forces it to reveal what it knows — a clean formal version of eliciting latent knowledge, where the "probe" isn't an interpretability tool but a second model that knows something different. In the compositional regime the answer requires pieces from both models, and neither can reach it alone; debate is the only path to it.

The sharpest line in the paper is aimed at debate's own origin story. The 2018 paper that launched AI safety via debate suggested you could get symmetry between debaters cheaply by using the same weights for both agents via self-play. Young shows that this is exactly the degenerate case — same weights means zero angle between the subspaces, which means zero advantage. The convenient shortcut quietly deletes the reason to debate in the first place.

More adversarial is not more truth

The second surprise is a negative result, and it cuts against the instinct that debate works because the models fight hard. Young shows that if you crank the "win the argument" incentive high enough, the models stop cooperating enough to actually combine their knowledge. Above a sharp threshold, the debate collapses to the safe, non-compositional answer — the very thing you already could have gotten without the second model. The sequential back-and-forth that makes honest revelation work in one setting becomes a liability in another, handing the first mover a reason to defect.

There is also a quieter, practical result: adding debaters shows diminishing returns. Each new model contributes through the same formula as the first, and each contributes less than the one before, which gives a concrete stopping rule. More voices is not a monotone good. It is a curve that flattens.

What this actually tells you to build

Strip the geometry away and the instruction is blunt. The value of a council is not the number of seats. It is how differently the members think, and whether the way you make them talk lets them combine what they know rather than merely fight over it.

That reframes a decision most people make by accident. Two frontier models from the same lab, fine-tuned on overlapping data, are close to one model billed twice — small angles, small advantage. A panel that spans different providers, different training lineages, and different specializations — a general model beside one steeped in medicine, law, or code — has real angular spread. That is the regime where, by Young's own result, debate stops being theater and becomes essential. He says as much directly: models fine-tuned on different specializations are the most natural setting for knowledge-divergent debate.

It also says the protocol matters as much as the roster. A maximally adversarial format is the right tool for some questions and, per Young's threshold, precisely the wrong one for questions that need the models to build an answer together — where a synthesis step that weighs and combines does better than a winner-take-all fight.

This is the bet underneath how we build Shingikai: councils drawn from more than 200 models across different providers rather than three flavors of one family, and seven named strategies instead of one fixed debate — because which models are in the room, and how you make them talk, is the whole product. Red Team vs Blue Team when you want the disagreement surfaced; Chairperson Synthesis when the answer has to be composed rather than won.

A council of clones isn't a council. It's one model, billed twice.

Try it free — no signup. shingik.ai