Gemini 3.5 Flash is the newer model. On the hardest reasoning, it scores below the older Gemini it sits next to.
On Humanity's Last Exam, the 3.1 Pro still posts 44.4% to Flash's 40.2%. On ARC-AGI-2, 77.1% to 72.1%. The newer, cheaper Flash got faster and stronger at coding and agentic work — and quietly gave back ground on the deep reasoning you'd actually convene a meeting for.
The downgrade hides inside the upgrade
That's the part a version number can't show you. "Newer" went up on some axes and down on others, and the one it went down on is the one that matters for the questions you can't afford to get wrong.
Google's answer is to wait. Gemini 3.5 Pro — unveiled at I/O on May 19, with a 2-million-token context window and a "Deep Think" mode — is slated for general availability this month to close exactly this gap. Fine. But you're running Flash now, today, on real decisions. And Flash is the model the rest of the stack is busy wiring in everywhere — Apple's Siri, Salesforce's Agentforce, even the chairman slot in Karpathy's open-source council. If it's your one model, you inherited the regression and nobody asked you.
Engineers already have a name for the general version of this: silent regressions. A model update changes behavior with no changelog entry, and prompts that worked yesterday quietly get worse. You don't find out from release notes. You find out from a user.
One model can't tell you it got worse
Models drift, and they always will. A reasoning budget gets dialed down to save cost. A "faster" release trades depth for latency. A safety patch shaves a few points off a benchmark nobody on your team is watching. The problem was never that models regress. The problem is that a single model is its own only witness.
Ask one model whether it's slipped and it answers with the same degraded judgment you're trying to audit. There's no baseline in the room. A council has one. When a member that used to land the same answer as the others starts drifting on reasoning-heavy questions, that delta is the alarm — the regression shows up as fresh disagreement instead of a silent wrong answer you ship.
One model has a confidence level. A council has a second opinion that changes when the first one slips.
You can't grade the test with the student who failed it
This is the same trap as letting one vendor be both the answer and the judge. If Flash produces the answer and Flash also checks it, a regression in Flash's reasoning corrupts the check too. Same blind spot, twice. The only honest audit comes from a model that didn't regress in the same direction at the same time — which, by definition, means more than one model in the room.
The delta is the signal
Run the same hard question past Claude, GPT-5, and Gemini. On routine work they'll mostly converge, and convergence is cheap reassurance. The interesting moment is the one where a single member breaks from the pack on exactly the kind of question it used to nail.
That break is not noise to average away. It's a model telling on itself. Last month it agreed with the others on multi-step reasoning; this week it doesn't. A council surfaces that the instant it happens. A solo model buries it under the same even, confident tone it uses for everything — including the answers it has quietly started getting wrong.
You don't have to wait for next month's fix
Google says Pro will close the regression. Maybe it will. But "wait for the next version" is a vendor's timeline, not yours, and the decision in front of you is dated this week.
A council doesn't wait. When one member is weak on a class of question right now, the synthesis layer down-weights it right now. Survivor narrows to the answers that hold up under challenge. Red Team vs. Blue Team forces the weak reasoning into the open instead of letting it pass on tone. Chairperson Synthesis decides what to do when the members split — which is the entire job once you accept that they will.
The model you picked is not a fixed thing
Here's the uncomfortable version. The model you standardized on is a moving target, its quality set by a release schedule you don't control and a changelog that won't mention the parts that got worse. Betting one model is betting that the target never moves the wrong way while you're not looking. The Gemini numbers say it already did.
A council doesn't stop models from regressing. It refuses to let the regression be invisible.
The next time an upgrade is quietly a downgrade, you'd rather hear it from another model than from your users. Try it free — no signup. shingik.ai