Someone rebuilt Sidney Lumet's 12 Angry Men as an AI benchmark. Twelve agents, each handed one juror's personality from the 1957 film, dropped into the same murder case and told to deliberate. Across eighteen runs, seventeen ended in a hung jury.
The film is two hours of one holdout turning eleven certain men around. In the AI version, that turn almost never happened.
The room that won't move
The paper is "12 Angry AI Agents" (arXiv:2605.01986), and it is careful to call itself exploratory: two models, three conditions, three replications each, eighteen runs total. Small. But the headline is hard to unsee. Seventeen of eighteen deliberations deadlocked, and the authors name the reason plainly: anchoring is the dominant failure mode. The models plant a flag on their first read of the case and spend the rest of the deliberation defending it.
That should bother anyone selling the idea that you can put a handful of models in a room and let them talk their way to a verdict. The whole romance of 12 Angry Men is persuasion — evidence surfaced, minds changed, a unanimous acquittal built one juror at a time. The paper's finding is that the film's central event, gradual minority-to-majority persuasion, is precisely the thing current models won't do on their own.
A hung jury is only a failure if you demanded unanimity
Here is where most people read the result wrong. They see "the council couldn't agree" and conclude the council doesn't work. That is the correct lesson only if agreement was the product.
It isn't. A deadlock between twelve models that each dug in is not noise to be cleaned up. It is information: this case does not resolve itself, and anyone who tells you it does is hiding the split. The mistake is architectural, not conceptual. If you ask twelve members to also be the judge, you have built a system with no one to break the tie, and the paper shows they will not break it for you.
That is the argument for a synthesis layer stated as an experimental result rather than a slogan. Call it Chairperson Synthesis: the members argue, and a separate step decides what to do with the disagreement they surface. When the room hangs, the synthesizer is not a convenience bolted on top. It is the load-bearing component. Seventeen of eighteen times, it is the only thing standing between you and a shrug.
The finding nobody was looking for
The second result is stranger, and it opens a question the field has barely touched. The two models did not behave the same way.
GPT-4o changed its vote a mean of 1.0 times per run, across every condition. Llama-4-Scout ranged from 2.0 vote changes in the baseline to 6.0 under an open-minded prompt, and it was the only model to reach a NOT_GUILTY verdict at all — once, in the no-initial-vote condition. Same instruction to keep an open mind. Llama internalized it. GPT-4o ignored it.
The authors' read: the intensity of RLHF alignment training, not model capability, is the primary determinant of deliberative flexibility. Their line is worth quoting. "Flexibility, not capability, tracks human deliberation."
Sit with what that implies for how you build a council. The instinct is to fill the room with the strongest models you can afford. But strength, in the sense that leaderboards measure it, is not the trait that lets an agent actually be moved by a good argument. A panel of heavily-aligned frontier models might be the least deliberative room you could assemble — twelve confident voices, each unwilling to yield, hanging every time. A roster that mixes alignment styles, including lighter-aligned open-weight models, might be the one that can still change its mind.
That is a concrete, testable, and mostly unclaimed argument for cross-provider diversity, and it runs on a different axis than the usual one. The usual case for a mixed roster is that different models know different things. This is a case that different models can hear differently. One is about coverage. The other is about whether the argument in the room lands on anyone.
Both failure modes are real
Honesty check, because it would be easy to overstate this. Other work points the opposite direction: models that fold too easily, conceding correct answers under peer pressure, agreeing more while knowing less. That is sycophancy, and it is a genuine failure mode. This paper found the reverse — models that won't budge when budging is the whole point.
Both are true. A council can fail by yielding too much and by yielding too little, and which one you get appears to depend on who is in the room and how the room is arranged: personas or not, an initial vote or not, the alignment profile of the members. The uncomfortable takeaway is that "just add more models" predicts neither outcome. Composition and structure do. A poorly built council does not average out its members' flaws. It picks one of two ways to fail and commits to it.
What the room is for
The clean version of the multi-agent pitch was always that a group of models talks itself to a better answer than any one of them would give. This paper is a quiet correction. Left to themselves, twelve models mostly re-elect their first guess and hang.
That does not make the council worthless. It makes the design of the council the entire game. The value was never that the members agree. It is that they surface where the real disagreement lives, and that a layer above them decides what to do with a split the members will not resolve on their own. Get that layer right and a hung jury stops being a dead end. It becomes the most honest thing the system can tell you: this one is not settled, and here is exactly where it comes apart.
One model hands you a verdict and never shows you the doubt. A council built correctly shows you the doubt first, then decides.
If you want to watch models actually argue a hard question and see where they split, that is the whole point of Shingikai. Try it free, no signup. shingik.ai