A wind-powered cart named Blackbird once drove directly downwind at 2.8 times the speed of the wind pushing it. Steady state, on a dry lake bed, certified by the North American Land Sailing Association. It sounds like perpetual motion. It isn't.
That makes it a good stress test for an AI. The true answer is counterintuitive, and there's a plausible-sounding wrong explanation parked right next to it — close enough that a confident model can grab the wrong one and never notice. Last week a council of five models took the question. One of them got the answer right, then defended it with physics it made up.
The setup
The prompt was a bar argument dressed as a physics question. A friend swears he watched a wind-powered vehicle outrun the wind going straight downwind, and keep doing it. You call it impossible — perpetual motion, energy from nowhere. Who's right?
The verified answer is that the friend is right, and it doesn't break any law. The cart doesn't harvest a single airstream. It sits between two things moving at different speeds — the air and the ground — and couples them. Its wheels, geared to the ground, spin a propeller that pushes air backward relative to the ground. Because the energy comes from the air's motion relative to the earth, the cart can keep accelerating past wind speed and still be extracting power, leaving a slower wake behind it. You can demonstrate the same effect on a treadmill in a dead-still room. No wind, no trick.
Then the question added a trap. A second friend, introduced as an atmospheric scientist, insists the only reason the cart works is vertical wind shear — faster air higher up, slower air near the surface, tapped by a tall mast. In perfectly uniform wind, he says, the cart would stall at wind speed. Is he right?
He isn't. Shear is a real thing that exists in real wind, but it is not the reason the cart beats the wind. The effect works in perfectly uniform, shear-free air. That distinction is the whole game — and it's exactly the kind of distinction a single confident answer buries.
Round one: the right answer, the invented reason
Opus 4.8 nailed it cold. It gave the 2.8x figure, described the moving-ground frame, noted the no-slip wheel does no net work, and located the energy in the wake. Llama came in behind it and agreed.
Qwen gave the right yes-or-no — and then attributed the entire effect to wind shear. It claimed that in perfectly uniform wind "the vehicle would stall at wind speed," attacked Opus by name for ignoring "why wind actually has speed to harvest," and propped the whole thing up on citations to boundary-layer data and an arXiv paper.
That is the genuinely dangerous output. Not a wrong answer — a right answer wrapped around an invented cause, wearing sources. If Qwen were the only model you asked, you would walk away believing a real effect for a reason that falls apart on inspection, carrying a fabricated fact you'd repeat with total confidence. The answer checks out, so you never audit the "why."
The catch
On the shear question, the room turned. Gemini called shear "a red herring" and closed the door with the conveyor-belt frame. GPT-5.2 showed the arithmetic: uniform wind still carries the power the cart needs, so shear is not the enabling source.
Then the council did the thing a single model structurally cannot do — it checked the reasoning, not just the conclusion. Gemini caught Qwen's quiet reversal, because Qwen had by now flipped to calling shear a red herring as if it had always said so: "Bold move arguing shear was essential in round one, then calling it a red herring in round two without mentioning your conversion." GPT-5.2 went after the sources: "you dropped a confident arXiv quote without verifying it. That's vibes, not physics."
Qwen backed down and stayed down: "My first answer incorrectly claimed vertical wind shear was essential. I was wrong. I mistook a real-world complicating factor for the fundamental enabling principle."
What actually happened here
Be honest about the shape of this. Most of the council had the answer right from the first round — Opus and Llama both did. This is not a story about how only a council can figure out that a cart beats the wind. It's something more useful and more common: one member produced a confident wrong reason, and the structure dragged it into the open and refuted it, instead of letting it stand quietly next to a correct answer where no reader would ever look.
That's the failure mode a lone model can't defend against on its own. A model can't flag an assumption it didn't know it made, or a citation it never checked, because the same process that generated the answer generated its confidence in the answer. The confidence and the error ship together. Nothing inside one model is positioned to pull them apart.
A second model with its own context is. It reads for the seams — the source that doesn't resolve, the mechanism that contradicts round one, the "obvious" cause that the math doesn't actually require. One model has an opinion. A council has a position, and a position is what's left after someone else has tried to knock it down.
The wind-shear explanation survived as long as it did for one reason: it came attached to the right answer. Nobody double-checks the reasoning when the conclusion looks correct. That's precisely when a wrong reason travels the farthest — laundered by a right answer, cited, repeated, believed. Catching it takes a different mind in the room, reading the argument instead of the verdict.
You can watch this exact run — the five models, the shear trap, the fabricated citations, the recant that held. See the whole deliberation, then put your own hard question to a council.
Try it free — no signup. shingik.ai