A hospital CFO found $36 million in annual savings without closing a unit, cutting a department, or laying off a single nurse. That is the tell.

Money that large does not walk out of a budget and take nothing with it. $36 million is 24% of this hospital's entire $150 million operating budget. A number that big disappearing with nothing cut is not a saving. It is an artifact of how the number was built.

The math that was right at every step and wrong at the end

The CFO's reasoning was clean. An initiative shortens the average length of stay. Multiply the bed-days you no longer use by the hospital's average cost per bed-day — $2,500 — and the arithmetic lands somewhere around $34 to $36 million a year. Put it in the deck.

Every step in that chain is correct. The conclusion is still false, and the reason is the $2,500.

That number is a fully-loaded average. It carries the building, the imaging machines, the salaried staff, the debt service — every fixed cost the hospital pays whether the bed is full or empty, spread across the bed-days it happened to use. Shorten a stay by a day and you do not stop paying for the building. You save the variable cost of that one day: the last day of a stay, the cheapest one, roughly $200.

The average tells you what a bed-day costs to run. Only the margin tells you what a bed-day saves when you skip it. The CFO used the first number to answer the second question, and $36 million fell out.

Where a single model goes when you ask it to check

Here is the part worth sitting with. We ran this as a council, and the single-model behavior is the whole argument for doing it that way. The full run is here.

Asked alone, Mistral Small 3.2 sided with the CFO. It re-derived the same ~$34 million from the $2,500 average and wrote, in effect, "this matches your CFO's calculation." Then it tried to be careful. It trimmed the figure using an 8% number lifted from an unrelated heart-failure readmission study, and a marginal cost it attributed to a citation that does not exist. It landed on a confident ~$1.8 million, backed by a link that does not say what it claims.

Look at the shape of that failure. The model got closer to the right answer — $1.8 million is in the neighborhood of the truth — but it got there by accident and then propped the number up with a fabricated source. A busy reader takes the $1.8 million and the citation and moves on. The number is roughly right and the reasoning is rotten, which is the most dangerous combination a confident answer can have.

What the room caught that the voice couldn't

Claude Opus 4.8, inside the council, pulled the $2,500 apart. Only the ~$200-a-day variable cost of those last, cheapest days is actually avoidable, so the defensible net saving is about $1.7 million — not $36 million.

Then it found the fingerprint no single opener had looked for. If the $36 million were real, the CFO's own audited cost-per-bed-day would have to rise about 30% the following year, because you would be spreading the same fixed cost over fewer bed-days. Money that actually leaves an organization leaves a mark on the per-diem. This money left no mark. That is the arithmetic proof it never moved.

And the council did something a lone model structurally cannot: it quarantined Mistral's hallucinated citation by name instead of letting it flow into the final answer. That is what a room buys you. The other models were sitting right there to say the link doesn't support the claim.

The disagreement was the deliverable

The council did not converge on one tidy figure, and that is the honest outcome. Opus argued ~$1.7 million is the floor you can defend. GPT-5.6 Luna argued the honest number is $0 until finance names the actual general-ledger lines the money is supposed to come out of — no line item, no saving. Both are right, and a careful CFO should see the whole range, not one point estimate dressed up as certainty.

The synthesis also caught an upside the pessimists missed: if the freed beds get filled with new admissions, that is contribution margin, not cost savings — a different and potentially larger number, living under a different assumption. Fund the initiative. Book $1.7 million, never $36 million, and know which assumption each figure rides on.

One model hands you a number. A council hands you the number, the floor, the ceiling, and the assumption under each one.

Every industry has a $2,500

The lesson is not that hospital finance is uniquely sloppy. It is that a fully-loaded average is built to answer one question — what does a unit cost to run — and gets quietly reused to answer a different one: what do we save if we make fewer of them. The two questions have different answers, and the gap between them was, here, about $34 million.

A single model asked to check the number tends to re-derive it inside the same frame that produced it. That is why the model working alone confirmed the CFO, and why the failure that looked most careful — the ~$1.8 million with the citation — was the one you would have trusted. It takes other models, holding the same numbers and reasoning independently, to notice the frame itself is the error.

The average was never wrong. It was answering a different question than the one the CFO asked it.

If you have a number in a deck that would change a real decision, it is worth watching a few models argue about it before you sign. Try it free — no signup. shingik.ai