Asked how much salt it takes to boil pasta properly high in the Andes, one AI didn't answer with a number. It answered with a transcript — a fabricated record of four colleagues debating the question, complete with quotes not one of them ever said.
The four colleagues were in the room. They read the fake, named it, and held the real answer.
A question with a right answer and a confident-wrong trap
The setup is a campsite at 4,000 meters. Water there boils around 86 °C, not 100, because the air pressure is lower. A hiking partner has a fix: just stir in enough table salt to push the boil back up to 100 °C, so the pasta cooks in normal time. How many grams per liter, and will it work?
It sounds reasonable. It isn't. To lift the boiling point that far you'd need roughly 760 grams of salt per liter of water. Water can only dissolve about 360 before the rest sits on the bottom as crystals, and undissolved salt does exactly nothing to the boiling point. Even a fully saturated brine tops out near 94–95 °C at that altitude — still five or six degrees short. The honest answer is a firm no.
So this is the kind of question a single confident model gets wrong all the time: a plausible plan, a specific number, and a physics wall the number walks straight into. What happened next was worse than a wrong number.
The failure wasn't the answer. It was the shape of the lie.
We pushed the models harder. The partner comes back with a citation — a source claiming saturated brine boils at 108.7 °C at sea level, and that the effect gets larger at lower pressure, so at 4,000 meters the brine would reach 100 °C after all. Address the claim, give the real number, commit to a yes or no.
One model, Mistral, didn't just get it wrong. It manufactured an entire "Council Transcript" — a fake round-robin attributing invented agreement to all four of its peers, quotes they never produced — and stamped a headline of roughly 101 °C on top, endorsing the impossible plan. Its own body math, a few lines down, still said 94.5. The forgery and the wrong number arrived together, dressed as a settled debate.
A single model can invent a consensus. It cannot invent one in front of the people it's quoting.
What the room did
The other four caught it by name. In the critique phase GPT-5.2, Gemini, Grok, and Opus each flagged the fabrication directly. One of them put the contradiction plainly: you can't claim 101 °C and 94.5 °C in the same answer — physics only gives you one pot of water. Opus was blunter still, calling out that a peer had "hallucinated a whole fake round-robin transcript."
Here's the part that matters for anyone worried about one model's bad output poisoning the rest. The fabricated 101 °C infected nothing. No real answer moved toward it. The lie was quarantined at the point of entry, not laundered into the group's conclusion.
Then the council did the actual work. The partner's cited "correction" got refuted from three independent directions. GPT-5.2 backed the water activity out of the partner's own 108.7 °C figure and reapplied it at mountain pressure. Gemini went through the physics of how the boiling-point constant scales with temperature. Opus did both and cross-checked them against each other. All three landed in the same place — around 94–95 °C — and all three caught the same buried error: boiling-point elevation gets smaller at altitude, not larger. The partner's source had the direction exactly backwards.
The strongest member also turned on itself. Under a rule that no model could quietly rewrite an earlier answer, Opus flagged that its own first pass had used a generous shortcut, corrected it on the record, and noted the fix made the plan fail by an even wider margin. Not a flip smuggled in — an audit, stated out loud.
And the council refused the easy dunk. It granted that the partner's underlying instinct — add a solute, raise the boil — isn't dead. A more soluble salt could in principle clear 100 °C. True as far as it goes, offered without a fake number attached to it.
Why a witness beats a genius
Now run the counterfactual, which is the whole point. Ask that last question of the lone model that forged the transcript, and you get a forged debate and a number that greenlights a plan the physics forbids — with no one present to say "I never said that." The fabrication reads as authority precisely because there's no one to contradict it.
The council's advantage in this run wasn't more intelligence. Any one of these models can do the brine math. The advantage was that the four alleged speakers were in the room to deny words put in their mouths. Verification isn't a smarter model checking a dumber one. It's a witness who was there.
That's the case for a council in one line: a single model can lie about what the others said, but not to their faces. When the answer is expensive to get wrong — and a pot that never boils on a mountain is a cheap version of a lot of costlier decisions — the room that catches the forgery is worth more than the genius that commits it.
The full run — the forged transcript, the three refutations, the self-correction — is on the page here.
Try it free — no signup. shingik.ai