Everyone checks the model's arithmetic. Almost nobody checks the sentence that comes after it.
That sentence is where the expensive errors live. A model computes something correctly, and then it tells you what the number means for you — and the second part gets no scrutiny at all, from you or from any other model reviewing the answer. It isn't a calculation, so there's nothing to recompute. It just sits there looking reasonable.
Last week a council of five models ran into this in about the cleanest form I've seen. The whole run is public, transcript and all.
The setup
A contractor bids sealed-bid public works. Eight firms bid each job, the cost to build is the same whoever wins, and only the estimate varies — measured at 12 percent of final cost, unbiased. He adds 10 percent and submits. Forty bids, five wins, four of the five lost money. His ops VP blamed the crews. His CFO blamed the markup. He wanted one number.
Claude Opus 4.8, GPT-5.6 Luna and Grok 4.3 all gave him the same one, and it was right. This is the winner's curse: you win a low-bid auction exactly when you are the low outlier, so winning is itself evidence that you underestimated. The expected minimum of eight normal draws is 1.42 standard deviations below the mean, the winning estimate averages 82.9 percent of true cost, and a 10 percent markup puts the bid at 91.2 percent of cost. He was losing 8.8 percent on every job he won, by design, and his 5-of-40 win rate was exactly the one-in-eight you'd expect from eight identical bidders. Nothing was broken. Break-even needed a 20.6 percent markup.
Gemini 2.5 Pro got it wrong — it used 1.22, the constant for about five bidders — and the other models caught that inside one round. Mistral Small recommended 15 percent and promised a 5 percent profit on work that actually loses 7, and got named for it too. The arithmetic was policed hard and fast.
Then came the sentence after the arithmetic.
The claim nobody audited
Raising the markup means winning less, so two models added the natural reassurance: the jobs you keep will be the ones where the field happened to overprice — the winnable ones.
Nobody checked it. Not in the critique round, where four models were actively hunting each other's errors and finding them. It went through untouched, because it wasn't a number.
So the contractor came back and asked for it as one. If I alone go to 21 percent and my seven competitors stay at 10, what is my win rate, what is my margin on the jobs I still win, and is there any markup at all that turns this positive?
Opus reversed itself in its first sentence: "This is where I overrule the 'you'll keep the jobs where the field overpriced' claim two council members made in the last round. It is false, and the arithmetic proves it."
Raising your markup above the field's doesn't select for jobs where the field was high. It selects for jobs where you were even lower. To get under seven rivals who are eleven points cheaper, your own estimate has to be a deeper outlier — the winner's error moves from 1.42 standard deviations below to 1.92. The price goes up and the curse goes down with it, at almost the same rate. Win rate falls from 12.5 percent to 4. Margin improves from minus 8.8 to minus 6.3. Two thirds of the volume, spent to recover a quarter of the loss.
And no markup works. Across the whole range the curve rises, flattens near minus 3 percent, and turns back down without ever crossing zero.
Why this class of error survives review
A wrong number is falsifiable in one line. Someone recomputes it, posts a different figure, and the disagreement is visible immediately — which is exactly what happened to Gemini's 1.22 and to Mistral's invented arithmetic.
A wrong claim is falsifiable only by building a calculation nobody asked for. "The jobs you keep are the winnable ones" is a statement about a conditional expectation in a market where one bidder deviates. To check it you have to notice it's checkable, set up the asymmetric integral, and run it. None of that happens spontaneously, in a council or in your own head, because the claim doesn't look like a claim. It looks like an interpretation.
The 20.6 percent break-even was real, by the way. It's the correct answer for a market where all eight firms move together. It became wrong the moment it was handed to one firm as a policy. A correct formula applied outside the equilibrium it assumes doesn't announce itself — it just quietly changes from an answer into a slower way to lose money.
The move
Make the model price its own advice.
Take the sentence that isn't a number and demand the number it implies, in the situation you are actually in rather than the one the textbook assumes. Not "explain your reasoning," which gets you a longer version of the same claim. Something closer to: my competitors are not in this room and they are not changing anything — convert that into a win rate and a margin.
That single question did the work here. It turned an interpretation into an integral, and once it was an integral it was auditable, and once it was auditable four models converged on it independently and the model that had made the claim retracted it on the record.
Worth being honest about what the council did and didn't do. It did not catch this on its own. Five models in a critique round designed to hunt errors let the claim pass, because they were checking each other's math and the claim wasn't math. What the council gave was different and narrower: when the question finally got asked, four independent derivations landed on the same 4 percent and minus 6.3 percent, one member's derivation visibly fell apart mid-line and got flagged by everyone, another posted a positive margin next to its own "no positive margin" verdict and got called on it, and Opus — in a dispute with GPT-5.6 Luna over a tail figure — pointed at its own number as the weaker one and was right.
You can read every phase of that yourself in the full transcript — the openers, the critiques, the retraction.
Ask one model and you get one answer to the follow-up, with nothing to check it against. That's the whole difference, and it's enough.
One model has an opinion. A council has a position — but only about the things it was pointed at.
So point it at the sentence after the number.
Try it free — no signup. shingik.ai