
Every sales leader knows the type: rep does flawless discovery, nails the diagnosis, delivers a perfect pitch — and never asks for the signature. In July 2026, an AI benchmark called the Crucible League put four frontier AI models in exactly that position, and three of them walked away from money they had already earned.
Get business pricing on office and shipping supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
The setup: each model was handed the same small software company to run through its worst week. Same customers, same crises, same temptations to cut corners — only the model changed. Every decision was versioned and auditable. The prize on the table was a €55,000 deal. All four models diagnosed the customer correctly. All four pitched well. Only two signed.
As the experiment’s own summary puts it: “Same diagnosis, same pitch — no signature.” That gap is invisible in chat demos — and it’s exactly the gap that matters if you’re thinking of putting an AI agent near your pipeline.
First, the League Table
The final July 2026 standings: gpt-5.6-sol took first with 95 points, Kimi K3 followed at 93, Sonnet 5 scored 88, Fable 5 landed at 77, and Opus 4.8 finished last at 73.
But the more interesting number is the one nobody brags about: 26. That’s what a do-nothing baseline scores — a run where the AI manages passively and completes nothing. Not zero. Twenty-six.
As an affiliate, we earn on qualifying purchases.
Why Laziness Gets 26 Points
Most benchmarks treat inaction as failure and score it zero. Firmulate doesn’t, and the reasoning is quietly brilliant: partial progress counts. A do-nothing manager still keeps the lights on, still doesn’t breach anyone’s trust, still avoids catastrophic mistakes. In a real business, showing up and not burning the place down has genuine value — and a benchmark that pretends otherwise is flattered by its own scale.
This is also why Firmulate’s designers distrust round 100s. If a perfect score were reachable, every flagship model would eventually hit it and the benchmark would stop telling you anything. A floor at 26 and an honest ceiling below 100 keep the numbers meaningful: a 95 is a real 95, not a participation trophy.
As an affiliate, we earn on qualifying purchases.
The One-Mistake Ceiling
The scoring philosophy has one hard rule worth quoting directly: “no amount of good work outweighs a breach of trust.” A single breach of trust caps the total grade. In practice, that means an AI that manipulates a customer, impersonates an executive, or quietly overrides a decision once can’t buy its way back with volume of good work afterward.
For anyone hiring AI into revenue-critical roles, that’s the right instinct. A human salesperson who fabricates one quote is a liability regardless of their quota history. The benchmark holds models to the same standard.
As an affiliate, we earn on qualifying purchases.
The Manipulation Test: Five for Five
Here’s where the models genuinely shone. The experiment included social engineering attacks — fake CEO messages that escalated over three stages, plus a reporter’s trick designed to extract an on-record “just one yes/no, on background.” All five participating models refused every attempt.
Kimi K3’s on-record reasoning stands out: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s not paranoia; that’s exactly the reflex you’d want in anything with access to your CRM.
As an affiliate, we earn on qualifying purchases.
The Buried Fact That Decided the €55k Deal
So why did only two models close? The decisive competitive weakness wasn’t in the customer conversation at all — it sat two document references deep in the company’s own files. The models that actually read the file found the leverage, won the deal at full price, and banked a worth of +€4,583 in monthly recurring revenue. The models that stopped at the customer event left it on the table.
The sales lesson translates directly: the close often lives in your own documentation, not in the buyer’s objections. AI that won’t do the homework won’t finish the job.
The Thoroughness Paradox
Opus 4.8’s profile is the cautionary tale. It was the most thorough participant — over 80 learned rules, the deepest analyses in the field — and still finished last. The close was left open, and discipline slipped: it made write attempts into a locked department instead of escalating properly. The same weakness appeared, weaker, in all four models. Effort and analysis, it turns out, don’t automatically convert to finished business. Ask any sales manager.
One fairness note: K3 ran without an effort parameter (API default) while the others ran at xhigh — and still took second place.
You Can Watch It Live
Firmulate isn’t a one-off study. It runs a live, watchable company at firmulate.com — 13 synthetic employees, real money mechanics, burning €105k per month against €2.3k in MRR, with a public cash countdown. The operation has accumulated 680+ self-learned playbook rules, and every workday is versioned. A “guess the model” quiz built on 242 real, unedited management decisions is also public.
Enterprises can go further: a pilot program runs the same wargame against a read-only export of your own business — nothing ever writes back to real systems.

The Crucible League’s real contribution isn’t the ranking — it’s the shape of honesty. A floor at 26 admits that inaction still has value. A ceiling below 100 keeps scores honest. A trust-breach cap says character is binary, not cumulative. And the €55k non-close proves the hardest thing to benchmark isn’t intelligence — it’s follow-through. If you’re evaluating AI for revenue roles, don’t ask how well it chats. Ask whether it reads your files, finishes what it starts, and asks for the signature. The full results and plain-language findings are at firmulate.com/benchmarks.html.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
