Most of the effort poured into AI debate goes into the mechanics. Who speaks first. Whether each model can see how confident the others are. How many rounds the argument runs before it stops. A new controlled study turned all of those dials, one at a time, and found they barely move the outcome.

What moved it was who was in the room.

The paper is arXiv:2511.07784, "Can LLM Agents Really Debate?", from Haolun Wu at McGill and Mila with Zhenkun Li and Lingyao Li at the University of South Florida. It asks a question the field keeps stepping around: when several models argue, are they actually reasoning together, or just taking a vote with extra steps? To answer it cleanly, the authors needed a task with a verifiable right answer, so they used the Knight-Knave-Spy logic puzzle — knights always tell the truth, knaves always lie, spies do either, and there is exactly one correct assignment. No judge model grading on vibes. Either the team got it right or it did not.

Debate helps. That was never the question.

First, the good news for anyone who thinks argument beats a single voice. It does, and not by a little. On a balanced team of mixed models, debate lifted accuracy on a size-4 puzzle from 17.33% to 69.33%. On a harder size-6 puzzle, 6.00% to 46.67%. On size-8, 3.33% to 36.00%. The gains shrink as the puzzles get harder, but the direction never flips. Letting models challenge each other beats letting one of them answer alone.

So the interesting question was never whether debate helps. It was what makes it help. The authors set up six factors and varied them independently: team size, composition, confidence visibility, debate order, debate depth, and task difficulty. Six knobs. Turn each, watch the accuracy move, see which ones matter.

Two of the knobs were not knobs at all

Here is the finding, in the authors' own words: "intrinsic reasoning strength and group diversity are the dominant drivers of debate success, while structural parameters such as order or confidence visibility offer limited gains."

Read that again with an eye on which factors landed on which side. The things that mattered — how strong the individual reasoners are, and how different they are from each other — are decided before the debate starts. The things that barely mattered — speaking order, whether confidence is shared, how deep the argument goes — are the procedural settings everyone loves to tune.

You can't tune your way to a good AI debate. You have to staff one.

Diversity did real work here, but only in company. The paper is careful about it: "diversity provides modest but consistent gains in stability and accuracy when strong reasoners are present." Difference alone is not the point. A room full of weak models that fail in different ways is still a room full of weak models. Strong reasoners who fail in different ways are the combination that pays. That is the case for a heterogeneous panel over five copies of one model, made from a logic puzzle instead of a pitch deck.

What the transcripts showed

The outcome numbers are half the paper. The other half comes from reading how the debates actually went, and two behavioral findings there are worth sitting with.

The first: "majority pressure suppresses agents' independent correction." When a wrong answer had numbers behind it, a model that privately had it right tended to fold. Not because it was out-argued — because it was outnumbered. This is the failure that vote-counting can't see, because the vote is the mechanism producing it.

The second is the flip side, and it is the whole game: "debate success depends on agents' ability to overturn incorrect consensus." The teams that won were the ones where a correct minority could drag the majority off a wrong answer instead of getting steamrolled by it. The value of a council is not that it agrees. It is that it can disagree its way out of a shared mistake.

The part that cuts against us, said plainly

There is a reading of this paper that is uncomfortable if you sell structured deliberation, and it would be dishonest to skip it. If speaking order and confidence visibility and debate depth give "limited gains," doesn't that undercut the claim that how you run the room is where the value lives?

It sharpens the claim rather than sinking it, and the distinction is worth being precise about. The knobs this study tested are shallow parameters inside one debate format. Turn-taking order. Whether a confidence score is displayed. How many rounds. Those are settings on a single protocol. They are not the same lever as choosing a different protocol entirely — a Red Team vs Blue Team setup where one side's job is to attack, or a Chairperson Synthesis where a dedicated model weighs the arguments instead of counting them. "Which seat speaks second" and "what is this room for" are different questions, and this paper only measured the first.

But the honest takeaway stands: you cannot fix a bad roster with a clever procedure. If the models are weak or interchangeable, no amount of turn-order tuning rescues the debate. Composition comes first. Everything else is downstream.

What to actually do with this

Get strong, genuinely different models in the room. That is the finding with the most evidence behind it, and it is the one most councils get wrong by stacking the same model under different names. Then put something on top that can act on the second finding — a synthesis layer that can recognize when a correct minority is being buried by majority pressure, and overturn the tally rather than ratify it. Weigh who is right. Don't count hands.

A controlled logic puzzle just told you the same thing the messy real questions do. The debate is only as good as the disagreement it can survive.

That is the bet Shingikai runs on: real models, chosen to fail differently, arguing where you can watch them. Try it free — no signup. shingik.ai