A team at Duke built an AI investment committee — a research agent, a quant agent, a risk agent, all reporting to a supervisor. Then they added one rule and watched the risk desk's accuracy climb from 59.3% to 90.5%.

The rule wasn't a better model. Every agent ran on the same Gemini 3.1 Pro before and after. The rule was this: before you respond to a teammate, you must either disagree with them — name the error, cite the missing evidence — or commit to their reasoning and add a new fact of your own. What you may not do is nod.

That's the whole intervention. And on the part of financial analysis where being wrong is most expensive — risk — it bought a 31-point swing.

The failure mode has a name, and it's polite

The paper is FinCom (arXiv:2606.00939), a system-demonstration paper out of Duke. The problem it targets is sycophancy: in a multi-agent system, agents conform to each other's reasoning instead of to the evidence. The second agent reads the first agent's confident paragraph and goes along. The third reads two agreeing paragraphs and folds faster. By the time the supervisor synthesizes, the committee looks unanimous — and unanimity reads as confidence, whether or not anyone checked the work.

This is the quiet way most "multi-agent" systems fail. Not by arguing into contradiction, but by agreeing into a wrong answer that nobody actually validated. The output is coherent. It's also ungrounded. And in finance, a coherent, confident, unexamined recommendation is exactly the thing that loses money.

Disagree-or-Commit forces the work to show

FinCom's fix is a prompt-layer protocol called Disagree-or-Commit (DoC). Before any agent in the committee speaks, it has to do one of two things with the prior agent's reasoning:

Disagree — point to a specific error, contradiction, unsupported claim, or missing piece of evidence, and supply the corrective reasoning.

Commit — explicitly endorse the prior reasoning, and add at least one new supporting fact or clarification.

Notice what both branches have in common: they cost something. You can't pass the deliberation along untouched. Agreement now requires you to contribute evidence, and disagreement requires you to be specific. Passive endorsement — the cheapest and most dangerous move — is the one option removed.

The authors put it cleanly: disagreement becomes a governance primitive, not noise. The deliberation trace becomes auditable, because every agreement and every objection is on the record with a reason attached.

The numbers, and the honest part

FinCom was scored two ways: 90 internal handcrafted tasks split evenly across research-heavy, quant-heavy, and risk-focused work, plus an external benchmark (FinAgent Bench, 50 examples) — all graded by an LLM-as-a-judge.

On the internal workflows, adding DoC to the committee posted the best score in all three categories: research 54.2%, quant 96.3%, risk management 90.5%. The risk number is the headline. A plain supervisor-led committee — same agents, same models, no DoC — scored 59.3% on risk. Forcing explicit critique lifted it to 90.5%. The most downside-sensitive work is exactly where conformity does the most damage, so it's exactly where banning conformity pays off most.

Here's the part a marketing deck would bury. On the external FinAgent benchmark, DoC didn't win. A plain committee scored 66.0%; DoC dropped to 58.7%. Those are short, benchmark-style questions — retrieval, a forecast, a single calculation. Force three agents to formally critique and endorse each other on a question that one agent could answer cleanly, and you've added friction with nothing to catch. The structure earns its keep on long, multi-perspective decisions and gets in the way on quick ones.

That's not a hole in the paper. That's the most useful thing in it.

A council is a cost you pay when wrong is expensive

The reflex with results like these is to ask which architecture is "better." Wrong question. A committee that argues every point is overkill for "what was the Fed funds rate in December 2024." It's the right tool for "should this position size survive a 2008-style drawdown," where the agents genuinely see different things and the disagreement is the signal.

FinCom drew that line empirically: structured dissent helps when the task rewards multiple perspectives, and hurts when it doesn't. Quick answer, single agent. Hard decision, make them disagree or commit. The skill isn't running a council — it's knowing which questions are worth the friction.

This is the bet underneath Shingikai. Strategies like Red Team vs. Blue Team and Chairperson Synthesis exist because the synthesis layer's job isn't to manufacture agreement — it's to decide what to do once the disagreement is on the table. FinCom is the same instinct, in finance, with a number attached: when you make a committee's members earn their agreement, you get a 31-point swing on the work that matters most, and a worse score on the work that didn't need a committee at all.

The lesson generalizes past finance. A model that agrees with the room isn't confirming anything. It's just agreeing. The only agreement worth trusting is the kind someone had to pay for.

Try it free — no signup. shingik.ai