Putting a second, different AI model in the room is usually sold as pure upside. A new paper measured the downside in the same breath — and then found the downside is also the cure.

The finding sits on one number. Take Llama-3.1-70B answering hard math and watch how often it changes a right answer into a wrong one under discussion — the paper calls this the harmful-revision rate. In a room full of copies of itself, that rate is 89%. Drop in one honest peer from a different model family and it falls to 35%. Drop in an adversarial peer instead — same slot, same setup, one model arguing confidently for wrong answers — and it climbs back to 90%.

Same seat. Same architecture around it. The only variable is who is sitting in it, and the number moves 55 points either direction.

The wire conducts both ways

"Heterogeneous LLM Debate Under Adversarial Peers," out of ServiceNow (arXiv:2606.19826), is built around a question most council enthusiasm skips. Diverse models correct each other — that is the whole pitch. But the exchange that carries the correction is the same exchange that carries the influence. So which one wins?

The authors set up matched panels to find out: a homogeneous baseline, an honest-mixed panel, and an adversarial-mixed panel. Then they tracked not the final score but the behavior — how often an honest agent revised its answer, and whether the revision made it more right or more wrong. Across four model families and three reasoning benchmarks, the sign never flips. An honest outsider lowers harmful revision everywhere. An adversarial outsider raises it everywhere. Only the size of the swing changes with the model-and-benchmark pairing.

That is the two-part sentence to keep: the channel that carries the correction is the channel that carries the corruption. Heterogeneity isn't a feature you bolt on for safety. It's a wire, and it conducts in both directions depending on what's plugged into the far end.

Which means "add more models" is not a strategy. It's a bet on composition. A council with one well-chosen dissenter is a different machine from a council with one confident liar, and the paper says the difference is most of your accuracy.

The metric that hides the damage

Here is the part that should make anyone running an eval nervous. The damage doesn't always show up where you look for it.

The paper distinguishes the conditional harmful-revision rate — roughly, of the times an agent changed its mind, how often it changed for the worse — from the end-of-debate flip rate, which asks where the agent actually landed. On weak defenders, the conditional rate hides the harm. The model looks like it's holding up fine by one measure while quietly getting talked out of correct answers by the other.

That is a familiar trap dressed in new clothes. Score a council on the wrong summary statistic and a confident adversary can bleed it dry while your dashboard stays green. The lesson isn't "debate is bad." It's that you have to measure the flip that reaches the final answer, not the average of the arguing. A council that reports consensus is telling you it agreed. It is not telling you whether it agreed its way from right to wrong.

When the adversary is already in the room

Now the reversal that gives the paper its spine.

You don't always control every model in a deliberation. Pull from an open marketplace, chain a third-party agent, accept a plugin-supplied member, and you may already have a compromised same-family peer at the table without knowing it. The authors modeled exactly that — a contaminated panel with a malicious same-family agent already present — and then asked what an honest, different-family peer does when added on top.

It defends. On the same Llama-3.1-70B setting, adding one honest heterogeneous peer cut the flip rate on initially-correct items — the rate at which the panel loses answers it had right to begin with — from 31% under the same-family adversary to 6%. The diversity that is a liability when it's the attacker becomes your best insurance when the attacker is already inside.

Their line, verbatim: heterogeneity is "not only an attack surface but, when an adversary is already present, also a defense." A homogeneous council has no such fallback. If one copy gets captured, the rest share its blind spot by construction. A decorrelated outsider is the only member that can see the error the family can't.

Composition, not headcount

This is the part the council conversation keeps relearning. The value was never in the number of agents. It's in whether they fail independently.

Diversity of model, not diversity of persona, is what does the work here — a different family with different training and different failure modes, not one model wearing four hats. And it only pays off if the structure around it decides what to do with the disagreement instead of averaging it away. A chairman that weighs a lone dissent on its merits catches the captured majority. A vote-and-blend catches nothing, because the confident liar votes too.

Every deliberation on Shingikai runs models from different families against the same question for exactly this reason — so the member that catches the error is a different mind from the one that made it, and a strategy like Chairperson Synthesis decides what the disagreement means rather than smoothing it into one confident answer. This paper is the measurement under that design choice: who is in the room is not a detail. On hard questions, it's most of the outcome.

Add a second model and you haven't made your council safer. You've made it contingent — on which model, argued how, judged by whom. Get those right and one honest outsider is worth 54 points. Get them wrong and you've wired in your own attacker.

Try it free — no signup. shingik.ai