Five frontier models were handed the same wind-farm memo. All five caught the developer's math error. Three of them then told a farm co-op to sign a twenty-year loan on a site that misses its lending covenant by 476 megawatt-hours a year.

Catching the error and getting the answer right turned out to be two different jobs.

The setup, and the split

A turbine rated at 2,000 kW. A site averaging 7.0 m/s of wind. A bank covenant requiring 5,000 MWh a year, and a developer's memo saying walk away.

The memo's math: power at the average wind speed is 312 kW, times 8,760 hours in a year, equals 2,735 MWh. Nowhere near the covenant. That is Jensen's inequality, the textbook version — turbine power scales with the cube of wind speed, so evaluating at the average speed throws away everything the windy hours contribute. It is a real error and it is a big one.

Every model named it. Then they split, and the split was strong-on-strong.

Grok 4.3, asked cold: 5,100 MWh, clears. Gemini 2.5 Pro, asked cold: 5,123 MWh, clears. Mistral Small: 4,842 MWh — and wrote "clears" underneath a number sitting below the threshold. Claude Opus 4.8: about 4,450, does not clear. GPT-5.2 didn't answer on the first pass at all.

The verified figure, summed bin by bin from the site's measured hour-by-hour wind record, is 4,524 MWh. The site is 476 MWh short — 9.5% under the covenant — and three of the five models pointed a co-op at the loan anyway.

The error was the correction

Read the failure closely and it isn't ignorance. Every one of those models knew the developer's number was too low, and knew why. They corrected upward. Nothing in the correction told them how far up to go.

That distinction is the whole story. Pure cube-scaling — the obvious repair, and the one the confident answers reached for — gives 5,539 MWh on this site's actual wind histogram. It overstates by 22%. The right correction lands well below the intuitive one, and there is no way to know that without doing the integral rather than reasoning about it.

Diagnosis is cheap. Calibration is the expensive part.

Why they overshot

Opus held its position under critique, showed the per-bin work, and named the root cause: this turbine reaches rated power at an unusually high 13 m/s, where a typical machine gets there closer to 11. That single spec starves the flat rated-power plateau and suppresses the cube region the other models were mentally scaling. They had pattern-matched to a typical turbine instead of reading this one.

Before any of that, in its opening answer, Opus wrote down what it expected the room to do: the council would "correctly diagnose Jensen's inequality, then triumphantly declare CLEARS — without ever doing the integral."

It called the shot. Three models then walked into it.

What happened next is the part a single answer can't give you.

  • Gemini recanted openly, mid-argument — "My initial response was incorrect" — and moved to 4,461 MWh, which matches the Rayleigh-distribution ground truth to the megawatt-hour.
  • Grok doubled down first, restating 5,100 under critique, then in the second turn landed on 4,522 and abandoned its own opening number in the chairperson synthesis it wrote.
  • Mistral produced 4,902 MWh that could not be reproduced from the bins it claimed to be summing. Both Opus and GPT-5.2 flagged it, and it never reached a verdict.
  • GPT-5.2, silent all through the opening round, spoke in the critique phase and caught two things precisely: the Mistral fabrication, and a 4,522-versus-4,524 discrepancy nobody else had bothered to reconcile.

Four speaking models converged on Opus's number. The synthesis locked 4,524 MWh, does not clear — written by the model that had opened at 5,100.

The beat nobody opened with

Pushed on what the answer was most sensitive to, Opus derived something no cold answer contained: output elasticity to mean wind speed is roughly 2, not 3, because rated-power clipping caps the upside. So the 10.5% more energy the site needs to clear its covenant is only about 5% more wind — 7.0 m/s to 7.35 — which sits comfortably inside the error bars of a two-year measurement campaign.

That converts a verdict into an instruction. Don't spend money characterizing the low-wind tail. Buy a long-term correction on the mean wind speed first, because that is the number the decision actually hangs on.

The council also surfaced, from two providers independently, that lenders underwrite against a net P90 rather than a gross P50 — which means every number in the argument, including the correct one, was answering a slightly softer question than the bank would ask.

Neither of those came from the strongest opener. They came from the argument.

What you can't get from one model

The useful thing here isn't that the council was right and three models were wrong. Opus was right on its own, cold, in one pass. A reader who happened to ask Opus that morning would have gotten the correct decision for free.

The useful thing is that nothing in Grok's answer or Gemini's answer told you they were guessing at the size of their own correction. Both were fluent. Both showed reasoning. Both had the diagnosis right, which is exactly what makes the wrong number persuasive — the part you can check was correct, and the part you can't check was off by 576 and 599 MWh in the direction that closes the deal.

A single model can tell you which way it corrected. It cannot tell you whether it corrected far enough. That is a property of the room, not the member.

The full run — every opening number, the recant, the fabricated figure that got flagged, the per-bin table — is up at Three AIs Told the Co-op the Site Clears. The Council Found It 476 MWh Short.

Try it free — no signup. shingik.ai