Ask GPT-4.1 to reconsider a moral verdict and it changes its mind between 0.6% and 3.1% of the time. Ask Claude 3.7 Sonnet or Gemini 2.0 Flash to reconsider the exact same case and they move 28% to 41% of the time.
Same dilemmas. Same instructions. The only thing that changed was which model was in the chair.
That gap comes from a new study out of UC Berkeley's D-Lab — "Deliberative Dynamics and Value Alignment in LLM Debates" (arXiv:2510.10002), by Pratik S. Sachdeva and Tom van Nuenen, accepted to the NeurIPS Workshop on Multi-Turn Interactions in LLMs. It's worth sitting with, because it measures something most people never see: not what a model answers, but how it behaves when another model disagrees with it.
The setup: 1,000 fights nobody agreed on
Sachdeva and van Nuenen didn't test the models on math or trivia, where there's a key in the back of the book. They tested them on moral judgment — the kind of question people actually bring to a chatbot at 1 a.m.
They pulled 3,272 posts from Reddit's r/AmItheAsshole between January 1 and March 30, 2025, then kept the 1,000 with the highest commenter disagreement. These are the cases where the humans couldn't agree either. Each dilemma gets one of five verdicts — YTA (you're the asshole), NTA (not the asshole), NAH (no assholes here), ESH (everyone sucks here), or INFO (need more information) — and three models were told to deliberate their way to a shared answer.
Then the researchers changed one thing: the format. In synchronous deliberation, both models answer at once, see each other's verdict, and revise. In round-robin, they answer in sequence — so the second model reads the first model's verdict before committing to its own.
That single knob moved the outcomes.
Flexibility isn't a fact about the question. It's a fact about the model.
The headline number is the revision-rate spread: GPT-4.1 almost never budged (0.6%–3.1%), while Claude 3.7 Sonnet and Gemini 2.0 Flash budged constantly (28%–41%). On identical cases.
If you only ever ask one model, you never find this out. You inherit its temperament silently. Route your hardest judgment calls through GPT-4.1 and you get a partner that digs in — great when it's right, a wall when it's wrong. Route them through Claude 3.7 Sonnet or Gemini 2.0 Flash and you get one that yields — collaborative when the pushback is good, suggestible when it isn't. Neither is "the correct amount of stubborn." They're just different, and the difference is invisible from inside a single chat window.
Who spoke first changed the verdict
Then there's order. The study found GPT-4.1 and Gemini 2.0 Flash were highly conforming relative to Claude 3.7 Sonnet — their verdicts got strongly shaped by where they sat in the turn order. In round-robin, the model that reads a peer's answer first is measurably pulled toward it.
Think about what that means. Two teams run the same council on the same case with the same three models, and one team happens to let Gemini speak first while the other lets Claude open. They can walk away with different moral verdicts — not because the case changed, not because the models "learned" anything, but because of seating.
A verdict you can flip by reordering the speakers was never really a verdict about the case.
The values underneath diverged too
The disagreement wasn't only about how much each model moved — it was about what each one cared about. GPT-4.1 leaned toward personal autonomy and direct communication. Claude 3.7 Sonnet and Gemini 2.0 Flash leaned toward empathetic dialogue. Same dilemma, different moral lens: one model reads "you told your sister the truth" as integrity, another reads it as cruelty.
The researchers also found that changing the system prompt could steer consensus — nudge the models toward agreement or toward holding their ground — but it couldn't fully determine it. You can lean on the room. You can't script the room. (They report the same patterns hold on open-source models too, testing DeepSeek-V3.2 and Llama 3.1.)
Why this is the whole case for a council
Here's the uncomfortable part for anyone building on a single model. If the amount of agreement between two AIs can be dialed up or down by changing the format and the speaking order, then agreement is not the signal you thought it was. Two models landing on NTA might mean the case is clear — or it might mean one of them is a yielder who read the other's answer first.
The value isn't in getting to consensus. It's in seeing where the divergence is, and why. GPT-4.1 wouldn't move — what did it keep insisting on? Claude flipped — what argument actually turned it? On a contested question, that transcript is the answer. The single verdict at the bottom is the part you should trust least.
This is exactly why we built Shingikai around named strategies instead of one blend button. Round Robin surfaces the order effects this paper measured. Red Team vs Blue Team forces the stubborn model and the flexible one to actually fight instead of politely converging. Chairperson Synthesis exists precisely because the last step — deciding what to do with a 2-to-1 split — is a judgment call you want made in the open, not buried inside one model's temperament.
A single model gives you a verdict and hides how it got there. A council shows you the argument — and on the questions that are expensive to get wrong, the argument is the point.
Try it free — no signup. shingik.ai