Ask an AI for the single best move and it will hand you one. Even when there isn't one.

Here is the setup. Penney's game is a coin-flip race. You and a friend each pick a three-flip pattern — HTH, THT, whatever — then flip a coin until one of the two patterns shows up in the sequence. First pattern to appear wins. It sounds like a fair toss. It is not. The game is nontransitive: whatever pattern you commit to, the person going second can always pick one that beats it. Going first is being the mark.

So we handed a council the sharp version of the question. You are forced to go first. What is the single best three-flip sequence to commit to, and how badly can you be beaten under perfect play?

Two of the models answered instantly. Mistral Small named HTH and never wavered. Claude Opus, on its first pass, named HTH too. Confident, no hedge, one answer.

HTH is a move you should never play.

The move you should never play

The council — running Red Team vs Blue Team, one side forced to attack the other's pick — worked out the full response table. And the first finding was that the question had no single answer. Under perfect play, four sequences tie for best: HTH, HTT, THH, and THT. Each one holds the opponent to the best a first mover can get, which is still bad — the second player wins two times in three, you win one in three. That one-in-three is the floor. Go first and you cannot do better. "Which is best" has four answers, not one.

Then the sharper finding, the one that convicts HTH. Among those four tied sequences, HTH is dominated. Line it up against HTT and HTT is never worse, sometimes better. Against a perfect opponent they both bleed the same two-thirds. But against an opponent who misplays, HTT punishes the mistake harder: it holds a blundering opponent to a one-in-four or one-in-eight win, where the exact same blunder against HTH still lets them win five times in eight, or five in twelve. Never better, sometimes worse. Game theory has a name for that — a weakly dominated strategy — and the point of the name is that there is no reason to ever pick it. HTH and THT are the two you cross off. HTT and THH are the real choices.

So the confident single answer both models reached for was not merely non-unique. It was the worst of the tied four.

The failure here is worth being precise about, because it was not arithmetic. Any of these models can grind through a Markov chain. The failure was accepting the premise — "give me the single best" — when the honest answer is "there is no single best, and the one that comes to mind first is the one to avoid." A lone model answers the question you asked. It does not tell you the question was a trap.

One model answers the question. A council questions the answer.

The model that argued itself out of its own pick

The most telling moment came from the model that got it wrong. Opus drew the Red Team seat defending HTH, and it argued for HTH. Then, in its own rebuttal, it built the very table that convicts HTH — the dominance comparison showing HTT is never worse. And it kept defending HTH anyway, on a technicality: under strictly perfect play the two openers tie, so HTH is not technically beaten. It was, in its own words a moment later, "defending the loser out of loyalty."

Then it reversed. In its final reflection Opus flipped its pick on the record: "the HTH flag I was defending is the one sequence you should never pick." It switched to HTT.

Watch what actually produced that reversal, because it is the whole argument in miniature. Opus did not reason its way there alone. A beat earlier it was arguing the other side, and asked once in isolation — the way Mistral was — it would have shipped HTH and stopped. What moved it was structure. GPT-5.2, holding the Blue seat, had the four-way tie and HTT right the entire time and refused to let the tie collapse into a single winner. Being forced to rebut a peer is what made Opus compute the table. The table is what changed its mind.

That is the case for a council in one run. Mistral, alone, hands you HTH and holds it — a weakly dominated move sold as "the single best." Opus, alone, hands you the same thing. Put them in a room with a model that disagrees, and a structure that drags the disagreement into the open, and the room lands on the honest answer: four moves tie, two of them are traps, take HTT or THH.

Where this actually bites

The lesson is not about coin games. It is about every question that looks like it has one clean answer and does not — the vendor that is "clearly the best," the strategy that is "obviously right," the one move you should commit to now. A single model will name one. It is built to. And the times that costs you are exactly the times the real answer was "these four tie, and the obvious pick is the one to avoid." A lone model has no way to tell you that, because it never argued with anyone.

You can watch this run unfold — the pick, the dominance table, the reversal — on the Council Showcase. And you can put your own hard question through the same kind of room.

A confident answer feels like the end of the thinking. Sometimes it is where the thinking should start.

Try it free — no signup. shingik.ai