A single model, prompted once, wrote a better answer than the structured AI council did. It cost one sixty-second of the tokens to do it. That is the headline result of a new paper that spent its whole length building the council.

And it is not a failure. It is the most useful thing the paper proves.

The number that should stop you

The paper is "From Debate to Deliberation: Structured Collective Reasoning with Typed Epistemic Acts" (arXiv:2603.11781), a single-author study by Sunil Prakash at the Indian School of Business. It introduces Deliberative Collective Intelligence — DCI — a tightly specified protocol for making multiple AI agents reason together. And then, unusually, it runs a negative control and reports the bad news at full volume.

On routine tasks, DCI scored 5.39. Plain unstructured debate scored 8.58. The deliberative machinery didn't just fail to help — it made the answer significantly worse than every baseline, a gap of 3.19 points with a confidence interval that never touches zero. On overall answer quality across the whole study, a single agent generating once beat the council outright. And DCI costs roughly 62 times the tokens of that single agent to lose that comparison.

Most multi-agent papers would bury those lines. This one builds its argument on top of them.

The product was never the answer

Here is the reframe the numbers force. A council's deliverable is not the answer — it is the record of what it almost decided instead, and why it didn't.

DCI guarantees that every session ends in a "decision packet": the selected option, the residual objections that were never resolved, a minority report from whoever dissented, and the conditions under which the decision should be reopened. That packet was complete 100% of the time. A minority report showed up in 98% of runs. For every baseline the paper tested, that same minority report appeared in at most 16% of runs — and on one table, in zero percent of them.

So the trade is stark and clean. A single model gives you a slightly better answer, far cheaper, with no idea of what it ruled out. DCI gives you a slightly worse answer, far more expensively, plus a full account of the disagreement underneath it. One model has an opinion. A council has a position — and a paper trail.

Where the structure actually earns its cost

The averages hide where DCI wins. On non-routine tasks, it beat unstructured debate by 0.95 points (95% CI [+0.41, +1.54]) — a real, significant gap. On hidden-profile problems, where the right answer only emerges if scattered clues held by different members get surfaced, DCI scored 9.56 — the highest score any system reached on any domain in the study, beating even a strong single agent.

That is not a coincidence about which domains are hard. It is a statement about which domains are deliberative. A hidden-profile task is one where no single member can be right alone, because no single member has all the evidence. Structure that forces each member to put its private information on the table is the whole game. On a routine task, there is nothing hidden, nothing to surface, and the same structure is just overhead — 62 times the overhead.

This maps cleanly onto a decision most people get backwards. The question is never "is a council better than a model." It is "is this the kind of question where being wrong is expensive and the evidence is scattered." When it is, the packet is worth it. When it isn't, ask one model and move on.

The caveat that points straight at the real lever

There is one more line in the paper worth holding onto. DCI ran its main experiments on copies of a single model — Gemini 2.5 Flash talking to itself. The author flags this directly and runs a small preliminary test with a genuinely mixed council, two Gemini agents and two GPT-4o agents. Architecture-design quality rose from 8.13 to 8.71.

That is a small sample, and the author says so. But it points at the lever the main study deliberately held fixed: the diversity of the members. A council of one model wearing four hats is not the same instrument as a council of four different models that actually see the problem differently. The paper measured what structure alone buys. It left the gains from heterogeneity mostly on the table — and got a hint they're larger.

The paper also names the failure mode that makes all of this matter: convergent-but-wrong. Structure can manufacture confidence. A process that always terminates, always produces a clean decision packet, can hand you a beautifully documented wrong answer and make it feel earned. The minority report is not decoration. It is the one part of the packet that argues with the conclusion — which is exactly why it's the part worth reading first.

What this means if you're choosing

DCI-CF, the algorithm underneath all this, is framed against Arrow's impossibility theorem — the proof that no aggregation rule can satisfy every fairness condition at once. DCI's answer is modest and honest: it doesn't promise the right decision, only a terminating, transparent one. Procedure, not truth. That's the correct ambition. A council can't guarantee it's right. It can guarantee it showed its work.

So the real question the paper leaves you with isn't whether to trust a council over a model. It's whether the decision in front of you deserves a minority report. For a reversible, low-stakes call, it doesn't — the single model and its missing paper trail are fine. For an expensive, contested, evidence-scattered one, the answer was never the point. The dissent is the deliverable.

That's the gap between asking and deliberating. The whole reason Shingikai runs a Chairperson Synthesis strategy — where the synthesis layer's job is to decide what to do with disagreement, not to erase it — is that on the questions that matter, the residual objection is the part you can't afford to never have seen.

Try it free — no signup. shingik.ai