Every benchmark for AI councils reports one number: did the group get more answers right than a single model. In a new paper out of the University of Maryland, that number went up. It is also the least interesting thing the paper found.
The study — Multiple LLM Agents Debate for Equitable Cultural Alignment, an ACL 2025 oral (arXiv:2505.24671) — measured something almost nobody measures. Not just whether debate makes a model more accurate. Whether it makes the accuracy gain land where the model was already weakest.
The number the average hides
The task was cultural etiquette. The team used NormAd-ETI, a benchmark of 2.6K short stories about social norms drawn from 75 countries. Give a model a scenario — how a guest should behave at a meal, when a gift is expected, what counts as rude — and ask whether the behavior fits that country's norms.
A single model does not fail at this evenly. It tends to know the norms of the countries that dominate its training data and fumble the ones that do not. Average its score across all 75 countries and you get a tidy headline. The headline says nothing about which countries it kept getting wrong.
That is the trap in almost every council result you have read. A higher average can mean real new coverage, or it can mean the same blind spots, better hidden behind the ones the model already handled. Parity is the question the average refuses to answer: did the council actually see the countries a lone model could not?
What debate did to the average, and to the tail
First the average, because it moved. The researchers ran seven open-weight models — LLaMA-3, Gemma-2, EXAONE-3, Yi-1.5, InternLM-2.5, Aya-23, and SeaLLM-3 — in 21 pairings, under two setups. In Debate-Only, two models argue toward a joint answer. In Self-Reflect+Debate, each model can choose on every turn whether to critique itself or engage the other.
Debate beat the single-model baseline in 19 of 21 pairings, for an average gain of 7.05% in accuracy. Across the board, mean accuracy went from 66.4% as single models to 76.3% after debate. Ten points, from making the models talk to each other rather than adding a bigger model.
Then the number that matters more. The best debating pairs did not just score higher on average — they scored higher on parity, the measure of how evenly the accuracy is spread across cultural groups. Gemma-2 paired with Aya-23 reached a parity score of 0.994. A Gemma-2 and EXAONE-3 pair reached 0.986. The single judge model they compared against scored 0.964. Higher is more equitable, and the councils were more equitable than the lone model — the gain concentrated where one model had been blind, not where it was already fluent.
Two small models reached a big model's answer
Here is the fact that should reorganize how you think about spending on this. Multi-agent debate let relatively small models — in the 7-to-9-billion-parameter range — reach accuracy comparable to a single 27-billion-parameter model.
Two small models that genuinely differ, made to argue, landed where one model three times their size landed alone. That is composition buying what everyone assumes only scale can buy. The council's coverage did not come from a smarter member. It came from two members with different holes in their knowledge, each catching what the other missed.
This is the whole case for a diverse roster in one experiment. Aya-23 was trained with heavy multilingual coverage. Gemma-2 and EXAONE-3 carry different strengths. Put them in the same room on a question about a country neither fully owns, and the disagreement is where the right answer surfaces.
The part that keeps it honest
Debate was not magic, and the paper is careful about it. Self-reflection alone — a model critiquing its own output with no second model in the room — improved accuracy by 3.26%. Real, but less than half of what debate delivered. A model checking its own work helps. A model checked by a different model helps more. Talking to yourself is not the same as being challenged.
And the pairing mattered enormously. When the researchers let an oracle pick the best model combination per question, accuracy over single models jumped 22.5% on average and as much as 41.7% for the strongest pair — EXAONE-3 with Aya-23. The lesson is not "any two models are better than one." It is that which two, on which question, is most of the game. That is strategy selection stated as a research finding: the council is only as good as the decision about who is in it and how they engage.
Why this is the sharpest council result of the year
Most of the multi-agent literature argues about the mean. Does debate raise accuracy, does it cost too much, does a single well-prompted model match it. Useful arguments, all of them measured on one axis.
This paper picked a second axis and showed the council wins on it too — and wins in the place that is hardest to fake. You can lift an average by getting better at what you were already good at. You cannot lift parity that way. Parity only moves when you cover the ground you were missing. The council covered it.
That maps exactly onto what a council is for. Not a better average. A blind spot one model cannot see in itself, seen by a second model that fails somewhere else. When we run a question with real cultural or contextual stakes, the value is not that the group agrees — it is that the disagreement drags in the perspective a single model would have skipped, and a synthesis layer decides what to do with it.
One model gives you an average. A council gives you a distribution — and the whole point is that someone finally looks at the tail.
Try it free — no signup. shingik.ai