Asked to set a 99% reserve for a small insurer, one AI committed to $537,515. To the dollar. And it called that number 99% safe.

For the company's real claim book, $537,515 would have covered about 64% of years. Not 99. A figure the board would have banked as a near-certainty was closer to a coin flip.

The number wasn't just wrong. It was precise.

The trap isn't that the model missed. It's that it missed and dressed the miss in six significant figures. "$537,515, 99% confident" reads like the output of a machine that checked its work. The precision buys trust the underlying estimate never earned.

Bank on it and the math is brutal. At that reserve the insurer has roughly a 36% chance of blowing through it in any given year — and about an 89% chance of at least one breach within five. Every breach is a bill the reserve was supposed to cover and didn't. That's the road to insolvency, walked confidently, on a number that looked like safety.

A reserve is a question about the worst year. The average can't see it.

A 99% reserve is a tail quantity. It's not about a normal year — it's about the worst year in a hundred. And the thing that hides the worst year is the average.

This insurer's book averages $500 a policy: 91% of policyholders file nothing, 8% file around $1,500, 1% file around $38,000. The mean is calm. The tail is not. A handful of large claims does all the damage, and the year-to-year spread is enormous — the per-policy variation runs almost eight times the mean.

You might think ten thousand policies would smooth that out. For a thin, symmetric risk, they would. For a book where a rare $38,000 claim dominates, they don't — the few big claims swamp the many empty ones, and the safety you expect from size never fully arrives. The average is a decoy. The reserve lives in the tail.

One model has a number. A council has a number it had to defend against four others — and that difference is the whole story here.

What five models did that one couldn't

The same question went to five models at once. They didn't cluster.

Grok said $650,000. Gemini said about $610,000. Mistral, about the same. GPT-5.2 said $537,515. Opus said $850,000. One identical question, a $312,000 spread. Each model had quietly smuggled in a different assumption about how heavy the tail is — and not one of them announced it. Asked alone, any of them would have handed you its figure with full confidence and none of the disagreement.

Then the disagreement went to work.

Gemini rejected GPT-5.2's low number outright — calling a light-tailed assumption "dangerously wrong" for a book this skewed — and pointed at the fix: the answer is in the insurer's own data, so measure the spread instead of guessing it. Opus reframed the problem. Stop averaging, it argued, and count: a 99% reserve is really a question about how many of the year's big claims land at once. It worked the tail directly and came out near $807,000.

The sharpest move wasn't the answer. It was the auditing. Opus caught that Gemini's own tidy normal-curve estimate — $778,902 — actually under-reserved, buying only about 98.5% coverage against a 99% target. It caught that Mistral had reported a "Monte Carlo simulation" returning $835,000 — a run that never happened, a fabricated result — and kept it out of the verdict. Google caught OpenAI. Anthropic caught Google, and caught the fabrication. The synthesis landed around $820,000, then did the thing a lone confident model almost never does: it flagged that even that figure is a floor, because the inputs themselves are noisy estimates. "The demand for one exact dollar figure is itself the risk."

Set the two outputs side by side. One model: $537,515, 99% confident. The council: about $804,000 to $820,000, plus the sentence that mattered most — the 99% label on the cheap number was really 64%. Same question. One handed the board a comfortable illusion. The other handed it a calibrated reserve and named the illusion out loud.

The reserve that fails you doesn't announce itself

This is what a Chairperson Synthesis is for — not to average five numbers into a sixth, but to decide what to do with the fact that they disagreed by $312,000. The disagreement was the alarm. A majority vote would have muffled it. A single model never rang it at all.

The reserve that fails you doesn't arrive looking dangerous. It arrives to six decimal places, sounding sure. The way you catch it is to make more than one mind answer, and watch where they split.

You can read the full run — every opening number, every catch, every retraction — here: the council on the 99% reserve.

Try it free — no signup. shingik.ai