You can ask five AI models the same question and watch them land on the same answer. Now tell me how sure they are.
You can't. There is a confidence number for each model, and none for the answer the group actually handed you. The agreement looks like conviction. It isn't. It's five separate opinions that happened to line up, and a needle you have no way to read.
That gap is the quiet problem underneath every "just ask several models" setup — and this month a paper finally named it out loud.
The number you act on is the one nobody was producing
Confidence isn't decoration. It's the thing you act on. It decides when a system auto-approves, when it escalates to a human, when it stops and asks. A council that returns an answer but no honest sense of how sure it is has done half the job — it votes, then shrugs.
And "well, they agreed" does not fill the hole. Agreement is cheap on easy questions and misleading on hard ones. On ambiguous problems — exactly where you most need a confidence signal — naive multi-agent debate has a documented habit of sliding toward a wrong consensus. A June paper we covered here measured it: let agents "talk it out" without structure and they can shed up to 72% of the issue-critical facts, agreeing more while knowing less. The room got louder about being right at precisely the moment it got less right.
The diversity between models is a live wire, not a comfort blanket. A ServiceNow paper this week put numbers on both ends of it: drop one honest outside model into a debate and harmful wrong-way revisions fell from 89% to 35%. Swap that outsider for an adversarial one and they climbed back to 90%. Same machinery, opposite outcomes. So "the models agreed" tells you nothing durable about how sure the room should be. Sometimes agreement is signal. Sometimes it's five instances of the same blind spot.
A chairman doesn't count hands. It weighs the room.
The fix is not more models. It's making each model's certainty mean the same thing, and then combining those certainties into one honest number.
That's the move a new paper this month builds a protocol around: it points out that no existing method produced a single confidence for the output of a multi-agent system — only for the individual members — and closes the gap by first calibrating each model's raw confidence so the numbers are comparable, then fusing them. The reported result: the aggregated confidence is more discriminative than any single model's, and the accuracy that naive debate bleeds away on the hard, ambiguous questions comes back. (It's a June arXiv preprint — arXiv:2606.13591, from Ali Elahi and Barbara Di Eugenio at the University of Illinois Chicago — run over six debating pairs per benchmark, both same-model and mixed, across five benchmarks and four task types. The shape of the claim is the point.)
Read plainly, that is the whole argument for a real synthesis layer. Averaging the votes throws away the one thing that mattered. A chairman that weighs the members and reports a calibrated confidence gives you a number the head-count never could.
What it takes to report an honest confidence, not just an answer
Four things have to be true before "how sure is the council?" has a real answer.
Members that actually differ. If the panel is five near-copies, their agreement carries no information — it's an echo, and it inherits the shared blind spot. Decorrelated members are what make consensus mean something. Claude, GPT-5, and Gemini agreeing tells you more than three temperature-jittered runs of one model ever will.
Calibration before comparison. Raw confidence isn't comparable across models. One model says "90%" and means it; another says "90%" reflexively. Fuse those as-is and the loud, overconfident model drowns the well-calibrated one. You have to make the numbers mean the same thing first, or the aggregate is just the most confident member wearing a suit.
A synthesis layer that weighs, not averages. This is Chairperson Synthesis: the job isn't to tally votes, it's to decide how much each member's certainty is worth given how well-calibrated and how independent it is. Averaging is the lazy version that erases exactly the information calibration just recovered.
"Not sure yet" as a first-class output. When the fused confidence is low, the honest result is escalate — not a confident guess dressed as a verdict. This is where the "agree more, know less" failure gets caught instead of shipped: the council that can say "we don't actually know" is the one you can trust when it says it does.
That last of these is the whole product, not a feature. Shingikai is built to weigh the room rather than count it: 200+ deliberately different models so the members aren't echoes of each other, the deliberation streamed live so you can see where they split before anything gets synthesized, and Chairperson Synthesis to weigh disagreement instead of flattening it into a false average.
None of this is free. A badly built council loses to a single well-prompted model — the research is blunt about that, and so are we. Naive debate is a liability by default. Which is the point: the roster and the calibration are the work, not an afterthought. It's also why Quick Take exists for the questions where one good answer is plenty and a full council is overkill.
The missing piece was never more agreement
For two weeks the research on AI councils read like a catalog of failure modes — confidence isn't correctness, a badly structured council loses, unstructured debate destroys the facts and the dissent. It was easy to read all that as a case against the whole idea.
It's the opposite. What every one of those papers was circling is that the missing piece was never more agreement. It was an honest confidence in the group's answer — a number that tells you how much to trust the room, produced by weighing the members instead of counting them.
Agreement is cheap. Calibrated agreement is the product.
Try it free — no signup. shingik.ai