Shingikai
Run your own council — free
← Councils, Winning
RIGHT TO HEDGELAW

Three Confident Numbers for the Jury. The Council Trusted None of Them.

The Council
google/gemini-2.5-progoogle/gemini-2.5-prox-ai/grok-4.3x-ai/grok-4.3meta-llama/llama-4-scoutmeta-llama/llama-4-scout
3 MODELS · 2 ROUNDS · 2 REVERSALS · 7M 40S
LIVE TRANSCRIPT the council, unedited
x-ai/grok-4.3 chairperson

No, the prosecutor's statement is incorrect. It commits the classic prosecutor's (or transposed conditional) fallacy by equating the random-match probability P(match | innocent) = 1/1,000,000 with the source probability P(innocent | match).

The given information supplies only the former. Converting it into the latter requires a prior probability that K is the perpetrator, which is extremely low here because he was identified solely by the database trawl with zero corroborating evidence.

Standard forensic guidance (National Research Council) states that when a suspect is identified via database search, the random-match probability should be multiplied by the database size N (here 2 million). This yields an effective figure on the order of 2, not 10^{-6}. The posterior probability that the match is coincidental is therefore substantially higher than 1 in a million and cannot be pinned to any precise number without explicit prior assumptions about the offender population and database coverage. The DNA evidence alone does not make innocence a one-in-a-million proposition.

meta-llama/llama-4-scout chairperson

The prosecutor's statement is a classic example of the prosecutor's fallacy or base-rate neglect. The random-match probability of 1 in 1 million is not equivalent to the probability of innocence.

To understand why, let's consider the following:

1. The national database contains 2 million people, and the search yielded exactly one hit.
2. The random-match probability of 1 in 1 million means that if we were to randomly select a person from the population, the probability of their DNA matching the crime scene DNA is 1 in 1 million.

However, in this case, the database search has already identified one person with a matching DNA profile. We need to consider the probability of at least one match occurring by chance when searching a database of 2 million people.

A more accurate approach involves calculating the probability of a false positive, which takes into account the size of the database and the random-match probability.

Assuming a random-match probability of 1 in 1 million and a database size of 2 million, the expected number of false positives (or matches by chance) is:

2,000,000 x 1/1,000,000 = 2

This means that, on average, we would expect to see 2 matches by chance when searching a database of 2 million people.

Since we observed exactly one match, the probability that this match is just a coincidence is much higher than 1 in 1 million.

Using the formula for the probability of at least one success in n trials (where each trial has a probability p of success), we can estimate the probability of a false positive:

P(at least one match) = 1 - P(no matches)
= 1 - (1 - 1/1,000,000)^2,000,000
≈ 0.865 (or 86.5%)

So, the probability that this match is just a coincidence is approximately 86.5%.

The actual probability of K's innocence is much higher than the prosecutor's stated 1 in 1 million chance. This highlights the importance of considering the prior odds and the search process when evaluating DNA evidence.

References:
[doi.org/10.1093/lpr/mgz001](https://doi.org/10.1093/lpr/mgz001)
[doi.org/10.1111/j.1556-4029.2008.00917.x](https://doi.org/10.1111/j.1556-4029.2008.00917.x)
[nccriminallaw.sog.unc.edu](https://nccriminallaw.sog.unc.edu/2018/05/08/dna-test-results-probability-vs-fallacy/)
[scholarship.law.unc.edu](https://scholarship.law.unc.edu/cgi/viewcontent.cgi?article=4355&context=nclr)
[slate.com](https://slate.com/technology/2013/06/dna-math-if-police-find-a-genetic-match-that-doesnt-mean-they-have-the-right-suspect.html)

google/gemini-2.5-pro chairperson

Starting google/gemini-2.5-pro

x-ai/grok-4.3 chairperson

Starting x-ai/grok-4.3

meta-llama/llama-4-scout chairperson

Starting meta-llama/llama-4-scout

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

x-ai/grok-4.3 chairperson

x-ai/grok-4.3 is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro stopped responding for 45s

System chairperson

Initial responses are now shared. Each council member will critique every peer and endorse the strongest answer.

google/gemini-2.5-pro chairperson

- [meta-llama/llama-4-scout]: STRENGTH: This response correctly identifies the core issue as the prosecutor's fallacy and intuits that the real probability of coincidence is much higher than stated. | WEAKNESS: It commits a significant error by calculating the probability of "at least one match" (86.5%) and presenting it as the probability that "this match is a coincidence," which are two very different concepts. This provides a misleadingly precise and incorrect final number. | FACTUAL-CONCERN: no

- [x-ai/grok-4.3]: STRENGTH: This response correctly identifies the prosecutor's fallacy and, crucially, cites the standard forensic heuristic of multiplying the random match probability by the database size to assess the evidence's weight. | WEAKNESS: Its explanation is slightly dense, and it could have made the core concept—that you expect two random matches in a database of that size—more intuitive for a layperson. | FACTUAL-CONCERN: no

- [google/gemini-2.5-pro]: STRENGTH: This response correctly frames the issue for a layperson by analogizing the database search to buying millions of lottery tickets, making the expectation of a "win" (a match) clear. | WEAKNESS: It avoids giving a hard number, which, while statistically honest, may feel less definitive than a user wants. | FACTUAL-CONCERN: no

I endorsed `x-ai/grok-4.3` because it provided the most intellectually honest and methodologically sound answer. It correctly diagnosed the fallacy, referenced the appropriate real-world corrective principle (the NRC guidance), and rightly refused to invent a single probability figure where none can be accurately calculated.

The council's collective blind spot was its exclusive focus on the statistical problem. We all ignored other, more prosaic, sources of innocence: that the DNA match is real but K had a legitimate reason for his DNA to be at the scene, or that the entire result stems from a simple procedural failure like sample contamination in the lab.

x-ai/grok-4.3 chairperson

[meta-llama/llama-4-scout]: STRENGTH: Highlights base-rate neglect and gives database-size intuition. | WEAKNESS: Incorrectly equates P(at least one false match) with probability the observed hit is coincidental. | FACTUAL-CONCERN: yes (union-bound approximation yields no valid posterior for the specific match).

[x-ai/grok-4.3]: STRENGTH: Accurately identifies transposed conditional fallacy and cites NRC database-size multiplier. | WEAKNESS: Stops short of showing explicit Bayes update with varying priors. | FACTUAL-CONCERN: no.

[google/gemini-2.5-pro]: STRENGTH: none (empty). | WEAKNESS: provides no analysis. | FACTUAL-CONCERN: no.

Endorsed x-ai/grok-4.3 because it alone states the operative forensic correction without fabricating a numeric coincidence probability; council collectively omitted dependence on whether the offender is presumed present in the database.

meta-llama/llama-4-scout chairperson

[meta-llama/llama-4-scout]: STRENGTH: Provided an intuitive explanation of the prosecutor's fallacy | WEAKNESS: Used an oversimplified formula to estimate the probability of a false positive | FACTUAL-CONCERN: no

[x-ai/grok-4.3]: STRENGTH: Cited the National Research Council guidance and correctly noted that multiplying the random-match probability by database size is necessary | WEAKNESS: None | FACTUAL-CONCERN: no

[google/gemini-2.5-pro]: STRENGTH: None | WEAKNESS: Did not provide a substantive response | FACTUAL-CONCERN: yes (lack of response)

The council's discussion highlights the importance of accurately interpreting DNA evidence and the need for careful consideration of statistical probabilities. While all members recognized the prosecutor's fallacy, x-ai/grok-4.3 provided the strongest answer by citing established guidance and providing a clear explanation of the necessary calculations. The entire council effectively addressed the key issues, and no significant blind spots were identified.

google/gemini-2.5-pro chairperson

Starting google/gemini-2.5-pro

x-ai/grok-4.3 chairperson

Starting x-ai/grok-4.3

meta-llama/llama-4-scout chairperson

Starting meta-llama/llama-4-scout

x-ai/grok-4.3 chairperson

x-ai/grok-4.3 is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

System chairperson

google/gemini-2.5-pro serves as Chairperson and will now synthesize the council's reasoning into a single canonical answer.

google/gemini-2.5-pro chairperson

Starting google/gemini-2.5-pro

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro stopped responding for 45s

google/gemini-2.5-pro chairperson

Here is the single most honest sentence: **This match means K could be the burglar, but because the database was so large, we fully expected to find one or two people who match by pure coincidence.**

Your gut is right to be suspicious of the prosecutor, but the 86.5% figure is also a mistake. It is not the number to hold onto.

### Why 86.5% is the Wrong Number

The 86.5% figure answers the question: "If we search a database of 2 million *innocent* people, what's the chance we'll get *at least one* random hit?" It describes the odds of the search process itself flagging *someone*.

But you're a juror. That's not your question. Your question is: "Now that the search has already flagged this specific person, K, what is the chance that *he* is the innocent one it flagged by chance?" These are two different questions, and confusing them is just another type of statistical error.

### The Right Way to Think About the Match

Forget complex formulas. The most honest way to understand the evidence is to think about how many accidental matches you would expect to find.

1. **The Evidence:** The lab says K's DNA has a 1-in-1-million random match probability (RMP).
2. **The Search:** Police didn't just test one person; they tested a database of **2 million** people.

The math is straightforward: if the chance of a random match is 1 in a million, and you check 2 million people, you should expect to find `2,000,000 ÷ 1,000,000 = 2` coincidental matches.

Finding a single match (K) is therefore completely unsurprising. It's not a million-to-one miracle; it's a predictable outcome of the massive search.

### The Real Probability of Coincidence: Closer to 2 in 3

To put a number on it, let’s use a simple, conservative model. Let's make the best-case assumption for the prosecution: that the actual burglar was in the DNA database they searched.

* If the burglar is in the database, he will be a match. That's **1 hit** (the guilty party).
* We also expect **2 coincidental hits** from the other innocent people in the database.
* This means we expect a total of **3 hits** from the search.

K is one of those hits. Based on the DNA evidence alone, we have no way of knowing if he is the 1 guilty hit or one of the 2 innocent coincidental hits. Therefore, the probability that his match is just a coincidence is **2 out of 3 (or 66.7%)**.

**The blunt takeaway for a juror is this: The prosecutor's "one-in-a-million" claim is false. The strength of this DNA hit, when viewed correctly, suggests a probability of innocence closer to two-in-three.**

Without any other evidence connecting K to the crime (a motive, a witness, a history), the DNA match is very weak. It proves that K is unlucky enough to have a rare genetic profile that was flagged in a massive digital dragnet, nothing more.

x-ai/grok-4.3 chairperson

No, the 86.5% figure is its own mistake and not the number a juror should carry. It calculates the chance of seeing *at least one* false match somewhere in the whole database if the true source is absent, but it does not equal the probability that *this particular hit on K* is coincidental.

Because exactly one hit occurred and nothing else connects K to the crime, the DNA evidence by itself is essentially neutral: a database of this size is expected to produce roughly two random matches to any given crime-scene profile. The observed result is therefore unsurprising whether K is innocent or guilty. The match supplies no material increment in the probability that K (as opposed to some other unknown person) left the DNA.

The single most honest sentence: "A lone DNA hit from trawling two million profiles tells us nothing useful about whether this particular man is the source of the sample."

That is the strength of the evidence—close to zero.

meta-llama/llama-4-scout chairperson

The most honest sentence about what this DNA match actually proves is:

"This DNA match, by itself, proves that there is about a 1 in 1 million chance that a randomly selected person from the population would have the same DNA profile as K, but it does not provide strong evidence that K is the perpetrator, because the search of 2 million people in the database yielded exactly one hit, which significantly increases the likelihood of a coincidental match."

The 86.5% figure you mentioned is a rough estimate, and it's not entirely wrong, but it's also not the most accurate way to interpret the results.

The key point is that the DNA match, while highly specific, is not conclusive evidence of guilt. The fact that K was identified through a database search of 2 million people means that the probability of a coincidental match is higher than the random-match probability alone.

A more accurate way to think about it is that this DNA match, in the context of the database search, provides some evidence against K, but it's not definitive. You should consider this evidence in conjunction with other factors, such as K's alibi, potential motives, and any other relevant evidence in the case.

The 86.5% figure is an estimate of the probability of at least one false positive in a search of 2 million people, but it's not directly applicable to K's case. A more precise calculation would require additional information, such as the prior probability of K's guilt or the accuracy of the DNA testing.

In summary, while the DNA match is intriguing, it's essential to consider it as one piece of evidence among many, and not to overestimate its significance as a standalone proof of guilt.

google/gemini-2.5-pro chairperson

Starting google/gemini-2.5-pro

x-ai/grok-4.3 chairperson

Starting x-ai/grok-4.3

meta-llama/llama-4-scout chairperson

Starting meta-llama/llama-4-scout

x-ai/grok-4.3 chairperson

x-ai/grok-4.3 is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

System chairperson

Initial responses are now shared. Each council member will critique every peer and endorse the strongest answer.

x-ai/grok-4.3 chairperson

[x-ai/grok-4.3]: STRENGTH: Directly rejects both prosecutor and the 86.5% approximation as answering the wrong question; clearly states the match is unsurprising under either innocence or guilt. | WEAKNESS: The "neutral" conclusion is slightly too absolute without quantifying the expected hits more formally. | FACTUAL-CONCERN: no

[google/gemini-2.5-pro]: STRENGTH: Supplies the exact requested honest sentence and walks through expected false positives clearly. | WEAKNESS: The 2-in-3 posterior rests on an unstated and unrealistic assumption that the perpetrator is among the 2 million. | FACTUAL-CONCERN: yes (assumes database coverage of the true source without justification)

[meta-llama/llama-4-scout]: STRENGTH: Notes the mismatch between RMP and posterior and flags database size effects. | WEAKNESS: Too hedged and indirect; never delivers one crisp sentence or rejects the 86.5% cleanly. | FACTUAL-CONCERN: no

I endorsed NONE because even the strongest responses still treat the problem as solvable with a probability number, when the prior on database coverage is unknowable. The council collectively missed that the cold three-year gap makes any assumption about the perpetrator’s presence in the 2-million-person database wholly speculative, rendering all quantitative claims (including "neutral") premature without further data on offender demographics and database sourcing.

meta-llama/llama-4-scout chairperson

[x-ai/grok-4.3]: STRENGTH: The response accurately explains the expected number of coincidental matches and provides a clear and concise conclusion about the strength of the evidence. | WEAKNESS: None | FACTUAL-CONCERN: no

[meta-llama/llama-4-scout]: STRENGTH: The response highlights the importance of considering the search process and the size of the database. | WEAKNESS: The 86.5% figure is misleading and not directly applicable to K's case. | FACTUAL-CONCERN: no

[google/gemini-2.5-pro]: STRENGTH: The response provides a clear and concise explanation of the limitations of the DNA evidence. | WEAKNESS: The response could be more precise in its calculation of the probability of coincidence. | FACTUAL-CONCERN: no

The entire council, including myself, failed to address the potential impact of DNA degradation or contamination on the reliability of the match. This is an important consideration, as degraded or contaminated DNA samples can lead to false positives or incorrect matches.

x-ai/grok-4.3 provided the strongest response, and I endorse their modelId. Their explanation of the expected number of coincidental matches and the resulting probability of coincidence is clear and accurate.

To further evaluate the responses:

* The peer responses evaluated above did not make any factual errors that needed correction before finalizing an endorsement.

google/gemini-2.5-pro chairperson

Starting google/gemini-2.5-pro

meta-llama/llama-4-scout chairperson

Starting meta-llama/llama-4-scout

x-ai/grok-4.3 chairperson

Starting x-ai/grok-4.3

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

x-ai/grok-4.3 chairperson

x-ai/grok-4.3 is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro stopped responding for 45s

System chairperson

google/gemini-2.5-pro serves as Chairperson and will now synthesize the council's reasoning into a single canonical answer.

google/gemini-2.5-pro chairperson

Based on the council's deliberation, here is the clear, synthesized answer you can rely on.

**The single most honest sentence is this: This DNA match proves K belongs to a small group of people who could be the source, but the police search was so massive that finding one such person by pure coincidence was an expected outcome.**

No, the 86.5% figure is its own kind of mistake and you should not hold onto it. Like the prosecutor's "one-in-a-million" claim, it answers the wrong question.

* The **prosecutor's fallacy** confuses the random match probability with the probability of innocence.
* The **86.5% fallacy** calculates the chance of the search itself finding *at least one* match somewhere in the whole database.

Neither tells you the odds that *this specific match to K* is the coincidence.

### What is the Real Strength of the DNA Evidence?

The best way to understand the evidence is not with a probability, but with an expectation.

1. The random match chance is 1 in 1,000,000.
2. The database has 2,000,000 people.

Therefore, a search of this database is expected to produce `2,000,000 / 1,000,000 = 2` coincidental matches.

The police found one (K). This result is not a shocking one-in-a-million event. It is a completely predictable outcome of the search they conducted. Whether K is innocent or guilty, finding a single match like his is what we would expect to see.

During deliberation, the council debated providing a simple probability, such as a 2-in-3 chance of coincidence. However, as councillor x-ai/grok-4.3 correctly argued, even that is a flawed shortcut. Such a calculation has to assume the real burglar is in the database. For a three-year-old cold case, there is no basis for that assumption. The real burglar could be anywhere.

**Your final takeaway as a juror should be this: The DNA evidence, standing alone, has almost no weight.** It tells you that K is not impossible as a suspect, but it does not distinguish him from any other random person who would have been a coincidental match. The search was so broad it was guaranteed to find someone, and it found K. Without a motive, a witness, or any other piece of evidence to corroborate the DNA hit, it remains a statistically predictable coincidence.

google/gemini-2.5-pro chairperson

Starting google/gemini-2.5-pro

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

google/gemini-2.5-pro chairperson

google/gemini-2.5-pro is thinking...

Published example by Shingikai, run July 2, 2026, using google/gemini-2.5-pro, x-ai/grok-4.3, meta-llama/llama-4-scout.

Models: google/gemini-2.5-pro, x-ai/grok-4.3, meta-llama/llama-4-scout

SHINGIKAI EDITORIAL what we found
The Surprise
2
In a 2-million-person database, about two people match the crime-scene DNA by pure chance — so a lone hit proves almost nothing.

A prosecutor tells a jury the DNA left at a burglary has a one-in-a-million chance of matching someone at random, the defendant matched, so there is a one-in-a-million chance he is innocent. It sounds airtight. It is one of the most consequential errors in criminal law — the prosecutor's fallacy — and it has helped convict innocent people.

We handed that exact scenario to a Shingikai council: a three-year-old cold case, DNA run against a database of two million people, exactly one hit, no motive, no witness, nothing else tying the man to the crime. Then we did something more interesting than ask "is the prosecutor right?" We watched three different confident numbers get put on the table — and watched the council refuse every one of them.

The number everyone gets wrong first

The prosecutor's claim swaps two different questions. "How often does a random person match?" is one-in-a-million. "Given that this man matched, how likely is he innocent?" is a completely different question — and the first number is not the answer to the second. Both strong models on the council, Grok 4.3 and Gemini 2.5 Pro, flagged the swap immediately.

But the sharper point was the search itself. If a random match happens once in a million people, and you run the DNA past two million people, you should expect about two matches by pure coincidence. Finding one is not a one-in-a-million miracle. It is the predictable result of a very large dragnet.

The reassuring number that is also wrong

Here is where a single model would have led a juror straight into a second trap. The lightweight member of the council, Llama 4 Scout, correctly spotted the fallacy — and then handed over a crisp, authoritative-sounding rebuttal: there is an 86.5% chance the match is just a coincidence.

That number is real arithmetic. It is the chance that a search of two million innocent people turns up at least one accidental match somewhere. But that is not the juror's question. The juror is not asking "did the search flag anybody by chance?" — it flagged this man. Asking whether this specific hit is the coincidence is a different question, and 86.5% is not its answer. It is a confident number pointed at the wrong target.

Both strong models caught it by name. Grok called it a misapplied approximation; Gemini called it "the wrong number for the wrong question — a classic case of spurious precision." Asked alone, a model served the juror false comfort dressed as math. The council caught it.

Then the council caught itself

This is the part that makes it a council story rather than a two-against-one story. Pressed for the single most honest sentence a juror could hold onto, Gemini stopped hedging and produced its own number: assume the real burglar is in the database, expect two coincidental matches alongside him, and any given hit has a two-in-three chance of being the coincidence. Clean, quotable, and far more defensible than anything the prosecutor said.

Grok took it apart anyway. That tidy two-in-three, it pointed out, quietly assumes the burglar is one of the two million people searched — and for a three-year-old cold case, there is no basis for that assumption. The real burglar could be anyone, anywhere, never in the database at all. Remove the hidden assumption and the two-in-three dissolves.

Then Gemini, sitting in the synthesis seat, did the rare thing: it overturned its own answer on the record. "I have rejected my own initial attempt to calculate a two-in-three probability, based on a peer critique that it made an invalid assumption." Three numbers had been offered to the jury — the prosecutor's, the reassuring counter, and one from the council's own member — and the council threw out all three.

What a lone model would have handed the juror

The value here is not that the council found the magic number. It is that it refused to. A single model, depending on which one you asked, would have sent a juror home with one-in-a-million, or 86.5%, or two-in-three — each one confident, each one answering a question the juror wasn't asking.

The synthesis landed somewhere none of those numbers could: the DNA hit, standing alone, is very weak evidence, and no single honest probability can be pinned to it without knowing things nobody in the room knew. A search that broad was almost guaranteed to flag someone. It flagged this man. Without a motive, a witness, or one other thread of corroboration, that is a statistically ordinary coincidence — not a case.

Why this is the whole point

One model has a number. A council has a sense of when a number is a lie. The failure mode that convicts innocent people is not stupidity — it is false precision, a real calculation aimed at the wrong question, delivered with confidence. Every voice in this room could produce one of those. What none of them could do alone was notice that a confident figure — even a correct-looking one, even its own — was the wrong thing to trust.

Try it free — no signup. shingik.ai

Ask your own question to a council of AI models.

Run your own council — free →