Put the same model in a council twice and it doesn't get wiser. It closes ranks.
That's the uncomfortable claim in a paper called "The Inverse-Wisdom Law" (arXiv:2604.27274). Fill a panel of AI agents with models from the same family — three instances of one brand, all reasoning the same way — and the paper argues that adding more critic agents makes the group more confidently wrong, not less. The extra voices don't stress-test the answer. They ratify it.
The naive council disappoints for a reason
This is the quiet reason "just ask a few models" so often lets you down. The pitch is seductive: run the question past several AIs, keep what they agree on. But if those several AIs are the same model wearing different hats, agreement isn't evidence. It's an echo — and the research says the echo hardens as you add participants.
A second paper sharpens the knife. "Debate or Vote" (arXiv:2508.17536), out of UW–Madison, took multi-agent debate apart to see which ingredient actually does the work: the agents' independent first answers, tallied as a majority vote, or the rounds of arguing that follow. Across seven benchmarks, the vote does almost all of it. In most cases, voting with no debate at all matched or beat full debate. The authors go further and model the debate as a martingale — a process whose expected outcome doesn't move. The arguing, by default, nets to zero.
You're reading it backwards
Put those two findings together and the naive picture collapses. A same-brand council closes ranks, and the debate that's supposed to save it is a coin-flip. Stop reading there and you'd conclude councils are theater.
You'd have it exactly backwards. Both papers are arguments for a council — just not the lazy version. They tell you precisely what a council has to be to earn its cost. Not more agents. Two decisions: who's in the room, and who synthesizes.
Who's in the room: diversity is the mechanism, not the decoration
The Inverse-Wisdom Law's real finding isn't "councils fail." It's that homogeneity fails, and it fails worse as it scales. Three GPT-5 instances don't correct each other because they share the same blind spots — the same training data, the same failure modes, the same confident wrong turns. Add a fourth of the same and you've added weight to the error, not a check on it.
The fix is decorrelation: models whose mistakes don't line up. Claude misses what Gemini catches. DeepSeek reasons its way to a different wrong answer than GPT-5 — which is exactly what makes the disagreement informative. A council of genuinely different models is a council whose errors cancel instead of compound.
Who synthesizes: the chairman sets the ceiling
The same paper argues something builders keep underrating — the synthesizer's identity gates the whole system. Pick a chairman from the same family as the panel and it can override a smarter dissenting critic, because it shares the panel's instincts about what a "reasonable" answer looks like. The council's integrity is lower-bounded by whoever synthesizes, no matter how good the critics are.
That reframes chairman selection from a default setting into the single most consequential seat in the room. It's not who tallies the votes. It's who decides what to do with the disagreement.
The structure that beats the martingale
"Debate or Vote" says the arguing nets to zero unless the protocol actively keeps the correct answer alive across rounds. That's the whole game. A naive debate lets the majority average the right minority answer away — the confident crowd talks the correct outlier out of it. The fix is structure that rewards finding the flaw instead of finding the consensus: an adversarial pass where one side's job is to break the answer, not bless it. Run that, and the debate stops being a coin-flip and starts paying a dividend. Skip it, and you've paid for a vote you could have taken in a single round.
This is the spine of a month of research, not one paper's hot take. Naive discussion erases up to 72% of the issue-critical facts (the "Deliberative Illusion" paper, arXiv:2606.03032). Averaging votes throws away how sure each model was. Forcing consensus propagates the error. And now: same-brand panels entrench it, and the vote — not the debate — does the work unless you build the structure. One instruction runs through all of it. Agreement is the failure mode a council resists, not the goal it chases.
Which models, and who decides
So the question was never "how many models." It was which models, and who decides. A council built from three of the same brand, tallying a vote, is the thing these papers just dismantled. A council built from deliberately different models, with a chairman chosen to weigh the argument rather than count the show of hands, is the thing they quietly endorse.
That's the version Shingikai runs by default — 200+ genuinely different models, a chairman you pick, and the deliberation streamed live so you can watch who held their answer and who caved. The disagreement is the product. The vote was never the point.
Try it free — no signup. shingik.ai