The most common way to build an AI council is to take one model and hand it three personalities. The optimist, the skeptic, the domain expert. Ask each in turn, collect the answers, take a vote. It looks like a debate.

Someone measured it. It reads more like one model talking to itself in three voices.

The measurement

A paper this spring ran the experiment directly. Three agents, all the same model — Qwen2.5-14B — separated only by role prompts, across 100 grade-school math questions. Then it measured how far apart the three "opinions" actually sat.

They sat almost on top of each other. Mean cosine similarity of 0.888. And when you ask how many genuinely independent directions three role-prompted agents span, the answer came back as an effective rank of 2.17 out of 3.0. Three costumes bought roughly two opinions. The paper calls it representational collapse: role conditioning shifts the surface phrasing but barely moves the underlying reasoning, so the agents crowd into a narrow cone and hand back near-duplicate evidence in different voices.

Carry the setting with the number — one model, one benchmark, one encoder, and the magnitude moves with how you measure it. But the direction is the point, and the direction is not subtle. The disagreement was in the wording, not the reasoning.

Why this isn't a one-off

This is the measured version of something the field has been circling for a month. Correlated members cap what any vote can recover — if the errors line up, aggregation has nothing to fix. A council built from models that already agree has little epistemic gain left to buy. And communication between agents tends to make them more alike over a conversation, not less.

Put those together and role prompts are the worst case, not a clever shortcut. They start from a single model — maximally correlated by construction — and then talk it into agreeing with itself faster.

Here is the reframe worth keeping: diversity is not a design intention you assert. It is a quantity you can measure. And in the most common council setup, it is smaller than the roster implies.

The wiring, not just the room

It gets worse when you look at how the members are connected rather than who they are. A separate line of work on open-ended idea generation — not factual accuracy, so read it as a signal about process, not a benchmark score — found that diversity collapse traces primarily to the interaction structure, not to the models themselves.

Dense communication topologies, where everyone hears everyone every round, accelerate premature convergence. An authority gradient in the roles suppresses dissent relative to flatter groups. And stronger, more heavily aligned models contribute diminishing marginal diversity even as their individual answers get better. The wiring quietly contracts each agent's range.

More talking is not more deliberating. Composition sets the ceiling. Topology decides how much of it you give back.

The uncomfortable part

The instinct, once you see the collapse, is to add rounds. Let them argue longer, force more exchanges, and surely the disagreement shakes loose.

The evidence points the other way. A third paper reports that much of debate's measured benefit comes from candidate generation — the fact that you sampled several answers at all — rather than from the debate itself. Plain majority voting over the same candidates does comparably. The talking was not where the value lived.

That is uncomfortable and it belongs in the post, because it kills the easy fix. If the gain was in having several genuinely different attempts, then more rounds of the same three costumes converging faster is the opposite of what you want.

What actually helps

The fix is not a longer conversation. It is different members and a structure that keeps them apart long enough to disagree.

Different models fail differently because they were built differently — different data, different training, different blind spots. Role prompts change the voice; different architectures change the reasoning. A council of genuinely different models starts with the correlation low instead of spending the whole debate trying to manufacture it.

Then the format matters more than it looks. A protocol that forces each member to engage the others' argument rather than restate its own — Round Robin, Red Team vs. Blue Team, Collaborative Editing — is a structural brake on premature convergence, not a stylistic flourish. The topology finding is what makes that claim empirical instead of aesthetic.

And when the members do turn out more alike than they looked, the last defense is the resolution rule. A synthesis layer that weighs the strength of an argument rather than counting the show of hands — Chairperson Synthesis — is the only thing that helps once you accept that the vote itself can be near-duplicate.

The point

Three role prompts is not three opinions. A council whose members would all fail the same question the same way is a monologue with a cast list, and no amount of additional rounds turns it back into a debate.

The fix is a thing you design, not a thing you assert: different members, and a structure that keeps them apart long enough to actually disagree. That is the whole difference between a council and a costume department.

Try it free — no signup. shingik.ai