Put five AI models on the same hard question and watch them converge. It feels like the system worked — separate minds, landing together, the answer earning its confidence from the agreement itself. A new paper out of Harvard and NTT Research says that feeling is sometimes worth exactly nothing. The models can agree when not one of them preferred the answer they agreed on.

That is not a rhetorical flourish. It is the finding.

The paper is arXiv:2603.24676, "When Is Collective Intelligence a Lottery? Multi-Agent Scaling Laws for Memetic Drift in LLMs," by Hidenori Tanaka, posted in March 2026. It asks the question the whole multi-agent field has been circling without naming: when an LLM population reaches consensus, does that outcome reflect collective reasoning, systematic bias, or mere chance? The question matters more every month, because councils, panels, and agent ensembles are being wired into decisions that used to belong to people — and the reflex is to read their agreement as a signal.

Tanaka's answer is that agreement can be an artifact of how the agents talk, not the truth of what they are talking about.

The mechanism is the uncomfortable part

To get underneath the behavior, Tanaka builds a minimal model he calls Quantized Simplex Gossip. In it, each agent keeps its own internal belief, but it does not share that belief. It shares a sample — one drawn answer — and its neighbors update on the sample. That one detail is the whole story. "One agent's arbitrary choice becomes the next agent's evidence and can compound toward agreement." By analogy with neutral evolution in biology, where a trait can sweep a population for no reason other than the luck of sampling, he calls the effect memetic drift.

Here is the part that should stop you. Even when no agent favors any label a priori — no bias, no better information, nothing — the population still breaks symmetry and reaches consensus. The agreement is real. The reasoning behind it does not exist. A vote counts opinions. Drift manufactures them.

Two ways to agree, neither of them reasoning

Tanaka's math predicts a crossover between two regimes. In the drift-dominated regime, "consensus is effectively a lottery" — the group lands somewhere, and where it lands is close to random. In the selection regime, weak biases get amplified and shape the outcome. Notice what is missing from both descriptions: a model working out the right answer and persuading the others with it. In one regime the council's answer is a coin flip wearing a confident face. In the other it is whatever small slant the roster happened to share, magnified by repetition.

The scaling laws are the practical warning. Tanaka derives how the drift depends on population size, communication bandwidth, and how quickly agents adapt to each other. Read plainly, that says the "more voices, more rounds" instinct — the one that makes people add a sixth model and a fourth round when they want more confidence — pumps more drift into the system, not more signal. You can scale your way to a more confident answer that knows less.

Set the two options side by side. A single model gives you one traceable chain of reasoning you can inspect and overrule. A naive council — one where every member reads the others and echoes what it sees — can hand you an answer carrying the confidence of five models and the information content of a coin toss. That is a worse position than using one model, because it feels safer and isn't.

What the paper implies about how to build a council

The honest reading is not "councils don't work." Tanaka is describing a specific failure of a specific design: agents that learn from each other's sampled outputs, in a setting with nothing to break the tie but noise. The mechanism points straight at the fix. If drift is fueled by members updating on one another's raw draws, then the design that starves it is one where members do not simply marinate in each other's outputs until they blur together.

Keep the members independent long enough to actually disagree. Route the disagreement forward instead of letting it dissolve into a majority. And put a layer on top whose job is to decide what a split means — not to report which answer got the most votes. That last piece is the one most council products skip, and it is exactly the component Tanaka's result makes load-bearing. When the members will drift toward each other if you let them, the thing that decides can't be the members. It has to sit above them.

There is a reason we let you watch a deliberation happen live rather than handing you a tidy verdict. A verdict hides whether the agreement was earned or drifted. The transcript shows you. When you can see a model change its answer because another model made a real argument — versus change it because the room was leaning — you can tell reasoning from gossip. A Chairperson Synthesis that reads that transcript and weighs the split is not a convenience feature. On Tanaka's account, it is the part that keeps consensus from being a lottery ticket.

The takeaway a practitioner can use tomorrow

If you are running any multi-model setup, the question to ask is no longer "did they agree?" It is "would they still agree if they hadn't been reading each other?" Agreement that survives independence is worth something. Agreement that only appears once the models start echoing is drift, and drift is indistinguishable from insight right up until it costs you.

Tanaka gave the field the math to name the difference. The design lesson is older than the math: a room full of people nodding is not the same as a room full of people convinced.

Agreement is cheap. Agreement you watched happen is the only kind worth trusting.

Try it free — no signup. shingik.ai