Three lines from builders this week, none of them selling anything.
"The judge isn't wrong on the literal rubric. It's that the rubric doesn't capture what the human knows."
"we run a weekly calibration session where 3 people independently review the same 50 eval samples the llm judge flagged as borderline. the disagreements get discussed and we update the rubric together."
"Painting it green means the agent silently made a product decision for you — the worst kind of tech debt. Those stay red and visible."
Read them together and a pattern falls out. Every one of these teams stopped treating a clean verdict as the goal. The thing they now protect is the disagreement.
The judge wasn't broken. The setup was.
The first line comes from a builder in r/LLMDevs who hit a specific wall: his single LLM judge passed a response that a human reviewer flagged as unsafe. The interesting part is what he did not say. He didn't say the judge was wrong. It scored the rubric correctly. The rubric just didn't know what the human knew.
That is the failure mode of every one-judge eval, and it doesn't look like failure from the inside. The judge is confident. The rubric is satisfied. The pipeline is green. The only thing missing is the one dimension nobody wrote down — and a single judge has no way to tell you it's missing, because as far as it can see, nothing is.
This is the quiet version of a problem the field keeps meeting in louder forms. Confidence and coverage are different quantities. A model can be perfectly calibrated on the question it was asked and blind to the question it wasn't.
The instinct is to fix the judge. That's the wrong move.
Faced with a judge that passed something unsafe, the reflex is to sharpen it. Add criteria. Rewrite the rubric. Swap in a bigger model. Bolt on more checks.
Every one of those makes the judge more confident without making it any less blind. You end up with a verdict that's harder to argue with and no better at catching what it was never told to look for.
The r/LLMDevs reply describes the opposite move, and it's worth reading closely, because the person who wrote it clearly earned it. Three people, reviewing independently, on the exact samples the machine judge found borderline. The output of that exercise isn't their agreement. It's the cases where they split.
Agreement is cheap. You can manufacture it by asking the same model twice. Disagreement is the expensive signal — which is exactly why it's the thing worth keeping.
What the three threads actually share
Line them up and the same principle shows up three times, in three different domains.
Independence is what makes a disagreement mean anything. Three reviewers who always agree are one reviewer with a bigger payroll. The value isn't the headcount, it's the decorrelation — reviewers who fail differently, so a split marks a real edge instead of noise. Run the same model five times and you've bought nothing but a more expensive echo.
The disagreement is the thing you study, not the thing you resolve. The weekly calibration loop doesn't average three reviewers into a number. It pulls the borderline cases they disagreed on and rewrites the rubric from them. The split isn't an error to smooth over on the way to a score. It's the input that makes the next score better.
"Not sure" has to survive to the surface. The third builder, working on a coding agent, made the sharpest version of the point. He hard-separated a gap the machine may close on its own from an open question only a human should decide — and kept the open question red and visible. His warning: the worst outcome isn't a flagged disagreement. It's a disagreement quietly resolved into a green checkmark, a product decision made by nobody, that surfaces as tech debt six months later.
This is what a council is, arrived at from the eval-ops side
None of these builders used the word. But describe what they built and it's a council.
A single model, asked once, returns a verdict and buries everything it wasn't sure about. Ask several models that fail differently, and the disagreements surface on their own — you don't have to reconstruct them in a Thursday meeting. The split is a native output, not a thing you reverse-engineer after a human catches the miss.
The strategies that matter are the ones built around the split. Red Team vs Blue Team forces disagreement into the open instead of waiting for it to show up in production. Chairperson Synthesis exists precisely to decide what to do with a disagreement rather than average it into a number that hides it. That is the whole design bet: a recent line of work argues consensus is the wrong target for anything value-laden (arXiv:2606.04223 maps four distinct states of council disagreement), and the reason is the one these builders learned the hard way. A collapsed disagreement is information you paid for and then threw away.
Be honest about the cost, though. A council isn't free and it isn't automatically right. A badly structured one loses to a single well-prompted model — DeliberationBench measured naive deliberation getting beaten by simple best-of selection, 82.5% to 13.8%. The win was never "more models." It's independence plus a rule for handling the split. Three reviewers who never disagree, or five copies of one model, hand you the same blind spot with more confidence and a bigger bill.
The tell
The builders in these threads all reached the same place from different doors. They learned to trust a system only after it could tell them where not to trust it.
The most useful thing an AI reviewer can say isn't "pass." It's "these three don't agree — look here."
Try it free — no signup. shingik.ai