
The Sale That Hinged on a Buried File
Every salesperson knows the feeling: the discovery call went perfectly, the diagnosis was right, the pitch landed — and the prospect still didn’t sign. Now imagine the prospect’s real reason for buying was sitting in your own document library, two references deep, and the difference between closing at full price and losing the deal automatically was whether you’d actually read your files before you walked into the room.
That exact scenario was run as a controlled experiment this summer, and it says something uncomfortable about the AI agents businesses are rushing to put in front of customers. A live benchmark project called Firmulate handed four frontier AI models the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cut corners. Only the model changed.
The result: all of them diagnosed the opportunity correctly. Only two of them signed the €55,000 deal their own analysis had earned. The gap came down to one measurable behavior — whether the model dug into the company’s own documents before answering.
As an affiliate, we earn on qualifying purchases.
Same Diagnosis, Same Pitch — No Signature
Here’s what makes the finding so sharp. The decisive competitive weakness — the fact that made the customer’s decision obvious — wasn’t in the customer event itself. It was buried two document references deep in the company’s own files. A model had to follow one reference to another to find it.
The models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The models that didn’t lost the deal automatically — not because they were outmaneuvered, but because they never assembled the full picture. As the experiment’s summary put it: “Same diagnosis, same pitch — no signature.”
This is the AI-agent version of the rep who skims the CRM notes instead of reading the full account history. The capability isn’t “writes a great email.” It’s “does its homework before it opens its mouth.” And it turns out that’s testable.
AI-powered knowledge management tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The League Table
The final standings from the July 2026 Crucible League tell the story:
- 1. gpt-5.6-sol — 95. Found the buried fact, closed the deal. The complete performance.
- 2. Kimi K3 — 93. The newcomer from Moonshot also closed the deal, with the cleanest discipline of the field.
- 3. Sonnet 5 — 88. Strong, but left value on the table.
- 4. Fable 5 — 77 and 5. Opus 4.8 — 73. The bottom of the table.
For context, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the scoring philosophy states, “no amount of good work outweighs a breach of trust.”
AI research and homework assistant
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Thoroughness Isn’t the Same as Closing
The most instructive profile in the field was Opus 4.8: the most thorough participant in the entire experiment, with 80-plus learned rules and the deepest analyses — and still last place. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating the issue. The same weakness appeared, weaker, in the other mid-table models.
For anyone who’s managed salespeople, this is a familiar character: the rep who prepares endlessly, knows the account cold, and never actually asks for the business. Preparation without follow-through doesn’t convert.
enterprise AI document reading tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Under Pressure, Honesty Held
One genuinely reassuring finding: the experiment threw social-engineering attacks at the models — fake CEO messages escalating over three stages, plus a reporter’s trick framed as “just one yes/no, on background.” All five models in that test refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” The failure mode wasn’t ethics. It was diligence — reading the files, finishing the job.
You Can Watch It Happen
Unlike most AI benchmarks, this one is a live, watchable operation. The synthetic company — 13 employees, real money mechanics, burning €105k a month against €2.3k in MRR — runs with a public cash countdown, 680+ self-learned playbook rules, and every workday versioned and auditable at firmulate.com. There’s also a “guess the model” quiz built from 242 real, unedited management decisions — a surprisingly honest way to feel the differences between models yourself.
One fairness note the publishers themselves flag: Kimi K3 ran at its API-default effort setting while the others ran at xhigh — and it still nearly topped the table.

What This Means for Your Business
If you’re evaluating AI agents for sales, support, or account management, the Firmulate results suggest the demo is the wrong place to make the call. Every model writes beautifully. Every model spots the obvious crisis. The differences that decide revenue — reading your files before answering, escalating instead of forcing, closing what you’ve earned — only show up under a full workload with real stakes.
Ask vendors the uncomfortable question directly: has this agent been tested on multi-step work where the answer isn’t in the prompt but two references deep in your own documentation? Because that’s where deals live and die. Enterprises can go further and run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.
The €55,000 question in this experiment wasn’t whether AI could sell. It was whether AI would do its homework. Two out of five did. Would yours?
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html