Run the same council on the same question twice and you can get two different answers. That is real. Ask three models to deliberate, read back the transcript, run it again an hour later, and the final call can land somewhere else.
Most people hear that and reach for the exit. If it isn't deterministic, the reasoning goes, it isn't trustworthy — and a coin flip with extra steps is still a coin flip.
That instinct is half right. The half it gets wrong is the whole point.
The part that actually matters
Here is where this bites. You are not asking a council what the capital of France is. You are asking it something expensive to get wrong — whether to sign the lease, which vendor to standardize on, how to sequence a launch you only get to run once. The kind of question that carries weight precisely because reversing it is costly.
For that class of question, an answer you can't reproduce is an answer you can't defend. "The council said so" is not a reason if the same council would have said the opposite on Tuesday. You don't just need an output. You need to know why it landed where it did, and whether that reasoning survives a second look.
So the instability is not a footnote. It is the first thing a serious person should worry about. The mistake is thinking the fix is to make the panel more stable.
Where the instability actually lives
A recent paper ran 750 debates between three-model committees and measured how much of the agreement was real. The results are unkind to anyone selling a raw panel as an answer machine. Change the prompt's tone from friendly to hostile and the reported full-agreement rate moves by 50.4 percentage points. Delete the instruction telling the models to push back, and dissent reverts toward agreement by 23.1 points. When the evaluation controlled for a reading-order bias, the "debate won" verdict — which showed up 66% of the time without the control — collapsed into 299 ties, with accuracy unchanged. First author Chen Qian; treat it as a preprint, not settled law.
Read that carefully. The agreement was moving with the framing, not the facts. Tone in, consensus out. That is not deliberation converging on truth. That is a room full of models being agreeable.
And you can't buy your way out with more seats. A separate benchmark study, DeliberationBench, put three naive deliberation protocols against a single well-prompted model on a weak five-model council. All three lost, roughly six to one, while burning up to 2.5 times the compute. More voices, worse answer, higher bill. A related result on judge panels found that nine correlated models supply only about two independent votes' worth of information, because they make the same mistakes on the same items. Headcount is not independence.
So the panel, by itself, is unstable — and stacking more models onto it does not rescue it. If the panel were the product, this would be the end of the story, and the story would be bad.
The panel was never the product
Here is the reframe. That instability is real and measurable, and it is exactly why the panel was never the thing worth paying for.
Reproducibility is not a property of the panel. It is a property of the layer that decides.
Think about what a raw panel actually hands you: a transcript. Claude, GPT-5, and Gemini talked, they diverged somewhere, they mostly agreed somewhere else, and now you are holding a wall of text and a vibe. Run it again, get a different wall of text. The variance you are seeing is the variance of an argument — and arguments are supposed to be a little different every time.
A decision is not. A decision reads the argument, finds where the members split and why, weighs the evidence behind each side, and produces one call with the reasoning attached. Do that well and the call is stable even when the conversation that fed it was not, because it is anchored to the strongest evidence rather than to whoever sounded most confident this run.
A panel is an argument. A decision has to be repeatable. The council isn't the product. The layer that decides is.
What that looks like when it is built on purpose
This is the part Shingikai is built around, and it is a design choice, not a feature list.
Chairperson Synthesis is that deciding layer. It does not tally votes and hand you the majority. It reads where the models split, why they split, and what evidence each side is standing on, then writes a single decision with that reasoning shown. Two runs can produce two different debates and still resolve to the same call, because the synthesis is bound to the evidence, not to the mood of the transcript.
Red Team vs Blue Team goes the other way and forces the disagreement to the surface on purpose, so the synthesis has something real to act on instead of letting a genuine split average quietly into a bland consensus. And because you watch the debate stream live, you can see exactly where it diverged and check whether the decision handled that split honestly. You are not trusting a black box that might have answered differently an hour ago. You are reading the seam.
That is the difference between a room full of opinions and something you can carry into a meeting.
The honest version
None of this makes the underlying models deterministic. They aren't, and pretending otherwise would be its own kind of dishonesty. The variance is real. What changes is what sits on top of it.
A panel of models is an argument. An argument is not a decision. The thing worth paying for is not the room full of voices — it is the layer that turns what they said into something you can act on today and still defend tomorrow.
Try it free — no signup. shingik.ai.