Every AI council on the market makes one silent decision before the models ever speak: it hands all of them the same briefing. Same question, same documents, same context window. A new paper out of Carnegie Mellon says that single default is the thing that quietly turns your council back into one model.

The line is blunt. "When all agents are given identical evidence, deliberation collapses into herding rather than genuine belief revision, leaving multi-agent systems little better than a single agent." Read that twice. The failure is not the roster. It is not the number of seats. It is that everyone in the room read the same page before the meeting.

The setup nobody thought was a setup

The paper is "Diverse Evidence, Better Forecasts: Multi-Agent Deliberation Under Information Asymmetry" (arXiv:2607.01661, a July preprint, not yet peer-reviewed). The authors go looking for the design lever that matters in multi-agent forecasting and land on the one choice every product makes without noticing it is a choice: what each agent is allowed to know.

Give five models the identical evidence and they tend to reason their way to the identical conclusion. That looks like agreement. It is actually redundancy wearing agreement's clothes. The authors measured it. Under fully shared evidence, deliberation drives inter-agent variance down 90 to 92 percent, which they call "near-complete convergence." The room isn't converging on the truth. It's converging on the briefing.

The fix is a briefing, not a better brain

Their move is to engineer the opposite of the default. They call it designed information asymmetry: split the evidence pool into a shared public subset and disjoint private subsets, so each model holds something the others can only get by actually talking. A single dial, the public ratio, controls how much everyone shares. Their system, InfoDelphi, routes the right evidence to the right agent, lets them trade written rationales instead of bare answers, and then weighs each model's confidence rather than counting hands.

The headline: InfoDelphi beats the strongest single-agent and multi-agent baselines "by 12 to 18 percent in Brier score and 4 to 8 percentage points in accuracy" on PolyGym, a benchmark of 375 binary forecasting questions pulled from real prediction markets. Brier score is an error measure, so lower is the win. InfoDelphi lands at 0.178 against 0.206 for a single model answering cold. The gap is the entire argument.

And here is the two-part result worth keeping. When the authors turn the dial to full sharing, which is exactly how a normal debate is run, performance falls back to 0.193. Their words: "too little shared context is as harmful as too much, but for different reasons. The former breaks communication while the latter eliminates diversity." Starve the agents of common ground and they can't talk. Feed them all the same thing and they have nothing to say to each other. The sweet spot is in the middle, and almost every shipping council is parked at the wrong end of it.

More deliberation is not more truth

The other finding cuts against the instinct to let the models argue longer. InfoDelphi peaks at two rounds. A third round makes it worse, 0.178 sliding back to 0.196. The explanation is clean: "two rounds suffice to propagate most private signals, and further deliberation leads to convergence toward a group consensus rather than continued refinement."

Translated: once the unique knowledge has changed hands, extra rounds don't dig deeper. They just pull everyone toward the center. The longer a council talks after the information has already moved, the more it sands off the disagreement that made it worth convening. A flip isn't a mind changed, and a consensus isn't a conclusion.

What this says about how councils actually get built

Step back and the paper joins a run of recent work all pointing at the same uncomfortable place. One Cambridge result proved that two models trained on the same corpus have zero debate advantage between them. Others have shown same-brand panels getting more confidently wrong as they grow. Now this one adds the inference-time version: even capable, distinct models herd when you brief them identically.

The common thread is that the lever everyone sells is not the lever that works. "200 models to choose from" is a menu, not a strategy. The thing that buys you a real second opinion is structural: different models, reading different things, kept in productive disagreement long enough to surface what each one saw and no longer than that. The roster is the easy part. The architecture around it is the product.

This is where the seven strategies stop being a feature list and start being the point. Red Team vs Blue Team exists precisely to manufacture the disagreement a herding council loses. Chairperson Synthesis is the confidence-weighted aggregation step this paper validates, the layer that decides what to do with a split instead of averaging it into mush. When we run a hard question through a council and the models land in different places, that gap is not noise to be voted away. It is the output.

One honest limit, because the paper states it and so should we. This is a forecasting benchmark with retrieved documents, tested on three models, GPT-5.4-mini, DeepSeek-V3.2, and Llama-4-Scout-17B. The asymmetry trick needs an evidence pool to partition in the first place, so it doesn't transfer unchanged to a question answered purely from what the models already carry. The mechanism is narrower than the slogan. The slogan is still right.

The authors close by saying the quiet part at full volume: removing information asymmetry "eliminates most deliberation gains," which makes "diversity of input the key enabler of effective multi-agent reasoning." The council was never the magic. The different vantage points were. If every seat read the same script, you didn't build a council. You bought one model and paid for five.

Try it free, no signup. shingik.ai