The whole appeal of watching AI models argue is that you can see who said what. Claude took one position, GPT-5 took the other, and the disagreement sits right there on the screen with names on it. A new paper out of Sharon Li's lab at the University of Wisconsin–Madison says the models can read those nametags too — and when they can, they don't argue harder. They defer.

That is the uncomfortable part, and it now has a coefficient attached to it.

The label you see is not always the label the model sees

"When Identity Skews Debate" (arXiv:2510.07517) starts from a distinction most council products have never had to make out loud. There is the label a human reads on the transcript — "Model A said X" — and there is whatever the next agent receives before it responds. Those can be the same label. Often they are. And when the second agent can tell which answer came from a peer and which came from itself, the paper says it stops weighing the argument and starts weighing the source.

The authors call the two failure directions by their right names. Sycophancy is uncritically adopting a peer's view. Self-bias is stubbornly adhering to your own prior output. Prior work poked at each one separately. This paper is, in its own words, "the first principled framework that joins sycophancy and self-bias to mitigate and quantify identity bias" in multi-agent debate. It formalizes a debate round as an identity-weighted Bayesian update — each agent revises its belief, but the revision is tilted by whose name is on the incoming answer.

The reason this matters for anyone running a named council is one line from the abstract: identity bias "is widespread, with sycophancy far more common than self-bias."

The danger isn't the model that won't budge

Read that asymmetry slowly, because it inverts the failure you would expect. A single model can't be sycophantic to a peer — it has no peer. A council can, and this paper says caving is the more common distortion, not the rarer one. The seat you have to worry about is not the one that digs in. It's the one that folds, and folds for the wrong reason — because a name it recognizes appeared next to the other answer.

That should bother every product in this category, including this one. Perplexity's Council Mode, Karpathy's llm-council, Shingikai — the shared selling point is that you get to see named models take real positions. If those same names are inside the prompt the members read, the disagreement you are admiring is partly a popularity contest.

A dial, not a hand-wave

The contribution that turns this from a worry into an engineering problem is the Identity Bias Coefficient (IBC) — "a principled bias metric that measures an agent's tendency to follow its peer versus itself." Positive IBC means the agent leans toward the peer: sycophancy. Negative means it leans toward its own prior: self-bias. For the first time you can point a number at a specific council seat and ask which way it is bending.

And once you can measure the distortion, you can test whether removing it does anything. Their fix is almost embarrassingly cheap: response anonymization. Strip the identity markers out of the prompt so an agent literally cannot tell "self" from "peer," which forces equal weight on identity. No retraining, no new model, no extra round.

The size of the effect is the part worth sitting with. On MMLU, Qwen-32B shows a Conformity–Obstinacy gap of 0.608 in the standard setting. Under anonymization that gap drops to 0.024 — what the paper flatly calls "a complete removal of identity-driven distortion." A dial that was near its extreme goes to roughly zero, and the only thing that changed was that the model no longer knew whose answer it was reading.

Two labels, two layers

Here is where the practitioner move lives, and it is a side-by-side, not a slogan. A named council and an anonymized council are two different instruments for two different jobs.

When the human is the audience, keep the names on. The whole value of watching Claude and GPT-5 disagree is that you can see who staked out what, weigh their track records, and overrule the one you distrust. Strip that and you have thrown away the transparency that made a council worth running in the first place.

When the members are the audience, take the names off. The agents don't need to know they're arguing with GPT-5 to evaluate GPT-5's argument. Telling them is how you leak brand loyalty into a process that was supposed to be about reasoning.

The mistake is assuming those two labels have to be the same label. They don't. The transcript you read and the prompt the members receive are separate layers, and this paper is a careful argument for treating them separately — show the names to the reader, hide them from the room.

That is not a research project. It's a council strategy: an anonymized-deliberation variant that sits next to Traditional Council and Red Team vs Blue Team as a legible option you pick when the question is one where deference would be expensive.

The lab keeps circling the same nerve

This is the same Wisconsin group behind "Debate or Vote" (NeurIPS 2025 Spotlight), which found that in a five-copy same-model panel, the majority vote does almost all the work and the debate rounds add little. Put the two results side by side and a program comes into focus. One paper says: when the arguing barely moves the answer, don't pretend the arguing is the value. This one says: when the arguing does move the answer, it's often moving it toward whoever spoke, not toward what was said.

Neither result is an argument against councils. Both are arguments against lazy councils — same model cloned five times, everybody's nametag in everybody's prompt, majority vote at the end. The lever was never a smarter model. It was what each model is allowed to know about the others.

The fix for a sycophantic council isn't a better member. It's a room where no one can tell whose idea they're about to agree with.

Try it free — no signup. shingik.ai