Every team building an AI council eventually reaches for the same lever: add more models. Nine opinions must beat three, and three must beat one. It feels like insurance — the more voices in the room, the less likely they are all wrong at once.
That intuition is wrong, and not for the reason most people assume. The problem isn't that the extra models cost money or slow things down. The problem is that most of those votes aren't votes at all.
The seats look independent. The errors aren't.
Here is the statistical fact the "add more seats" reflex ignores: a panel is only as strong as the independence of its members. If three models were trained on overlapping data, share an architecture, or come from the same provider, they tend to be wrong in the same direction at the same time. When one hallucinates a citation, the others nod along, because the blind spot is baked into all of them. Stack ten of those and you have not stacked ten opinions. You have stacked one opinion, ten times.
A recent empirical study across a large set of models found the uncomfortable version of this: the strongest, most accurate models tend to fail most alike. Capability and correlation rise together. So the panel you would most want to trust — all frontier models, all sharp — is exactly the panel whose errors are most tightly coupled. Its confidence goes up while its independence goes down.
The honest way to count a council is not by seats. It is by effective votes — how many genuinely independent signals you actually have. A nine-member panel of near-clones can be worth just a fraction of its headcount in effective votes, because most of them are echoes. The headcount on the box tells you nothing about that number.
Even when the models differ, agreement can be theater
Correlation is one way a council fakes consensus. Conformity is the other, and it is sneakier.
One benchmark put a weak five-model council through three of the standard deliberation protocols and had all three lose to a baseline that simply picked the best single model for the job — losing by roughly six to one, while burning about two and a half times the compute. More deliberation, worse answers, higher bill. The raw panel didn't aggregate wisdom. It aggregated noise.
A separate set of 750 debates between three-model committees, from a team led by researcher Chen Qian, showed why. Change the prompt's tone from friendly to hostile and the rate of full agreement moved 50.4 points. Delete the instruction telling the models to argue and the dissent quietly reverted — a 23.1-point swing — back toward going along. Strip a reading-order bias out of the judge and 66% of the "the debate changed my mind" verdicts collapsed into ties: 299 of them, with accuracy unchanged. The models weren't reasoning their way to consensus. They were responding to social pressure and formatting.
Newer work pushes the knife in further. Models that cave to a wrong majority in a debate often still hold their original, correct answer internally — they change what they say, not what they believe. The dissent is still in there, suppressed. A panel reporting unanimous agreement can be sitting on live disagreement it pressured its own members into hiding. And other recent work keeps finding that structured debate beats flat majority voting as an aggregation rule, which only makes sense if counting raised hands was the wrong instrument all along.
Put those together and you get the real failure mode. A council can report confident agreement for two completely different bad reasons: the members shared a blind spot, or the members conformed. Flat majority voting can't tell those apart from genuine corroboration. It just counts heads.
Independence is the product. Headcount is the packaging.
So the design question is not "how many models?" It is two different questions that the seat count quietly conflates.
First: are the seats actually independent? Different providers, different training lineages, different failure modes — chosen so that when they agree, the agreement means something. A council you assemble to decorrelate is worth more at three seats than a provider-matched set is at nine.
Second: does the deciding layer read divergence, or does it count votes? A resolution rule that weighs where the models genuinely split — and treats agreement-among-clones or agreement-under-pressure as the weak signal it is — extracts more from three independent seats than majority voting extracts from twenty. The rule that resolves disagreement is doing the real work. The seats are just raw material.
This is the part the "put all the models in a room" products tend to skip. A fixed, often provider-clustered panel with flat majority synthesis assumes an independence it does not have, and then reports its highest confidence at exactly the moment every member is sharing the same mistake. The architecture is optimized for the number on the marketing page, not for the number that matters.
How we think about it at Shingikai
This is the whole reason Shingikai runs on 200+ models through OpenRouter rather than a house panel. The point of that breadth is not variety for its own sake — it is that you can draw seats to decorrelate, picking models that fail differently instead of models that fail together.
It is also why the strategies aren't interchangeable. Red Team vs Blue Team exists to manufacture independence when the models would otherwise agree too easily — you assign them opposing jobs so the dissent can't quietly revert. Chairperson Synthesis exists so the deciding layer can weigh divergence instead of tallying votes. And because you watch the debate stream live, you can see which kind of agreement you got: independent models converging on the same answer for different reasons, or three models nodding along to the first one that spoke.
That distinction — corroboration versus clones agreeing — is invisible on a scoreboard that only shows the final tally. It is the entire difference between a council that caught something and a council that just felt reassuring.
More models is not more reliability. Independent models is. A council of clones is one model in a trench coat.
Try it free — no signup. shingik.ai.