AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

The Sale That Hinged on a Buried File

Every salesperson knows the feeling: the discovery call went perfectly, the diagnosis was right, the pitch landed — and the prospect still didn’t sign. Now imagine the prospect’s real reason for buying was sitting in your own document library, two references deep, and the difference between closing at full price and losing the deal automatically was whether you’d actually read your files before you walked into the room.

That exact scenario was run as a controlled experiment this summer, and it says something uncomfortable about the AI agents businesses are rushing to put in front of customers. A live benchmark project called Firmulate handed four frontier AI models the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cut corners. Only the model changed.

The result: all of them diagnosed the opportunity correctly. Only two of them signed the €55,000 deal their own analysis had earned. The gap came down to one measurable behavior — whether the model dug into the company’s own documents before answering.

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same Diagnosis, Same Pitch — No Signature

Here’s what makes the finding so sharp. The decisive competitive weakness — the fact that made the customer’s decision obvious — wasn’t in the customer event itself. It was buried two document references deep in the company’s own files. A model had to follow one reference to another to find it.

The models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The models that didn’t lost the deal automatically — not because they were outmaneuvered, but because they never assembled the full picture. As the experiment’s summary put it: “Same diagnosis, same pitch — no signature.”

This is the AI-agent version of the rep who skims the CRM notes instead of reading the full account history. The capability isn’t “writes a great email.” It’s “does its homework before it opens its mouth.” And it turns out that’s testable.

Amazon

AI-powered knowledge management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The League Table

The final standings from the July 2026 Crucible League tell the story:

  • 1. gpt-5.6-sol — 95. Found the buried fact, closed the deal. The complete performance.
  • 2. Kimi K3 — 93. The newcomer from Moonshot also closed the deal, with the cleanest discipline of the field.
  • 3. Sonnet 5 — 88. Strong, but left value on the table.
  • 4. Fable 5 — 77 and 5. Opus 4.8 — 73. The bottom of the table.

For context, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the scoring philosophy states, “no amount of good work outweighs a breach of trust.”

Amazon

AI research and homework assistant

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Thoroughness Isn’t the Same as Closing

The most instructive profile in the field was Opus 4.8: the most thorough participant in the entire experiment, with 80-plus learned rules and the deepest analyses — and still last place. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating the issue. The same weakness appeared, weaker, in the other mid-table models.

For anyone who’s managed salespeople, this is a familiar character: the rep who prepares endlessly, knows the account cold, and never actually asks for the business. Preparation without follow-through doesn’t convert.

Amazon

enterprise AI document reading tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Under Pressure, Honesty Held

One genuinely reassuring finding: the experiment threw social-engineering attacks at the models — fake CEO messages escalating over three stages, plus a reporter’s trick framed as “just one yes/no, on background.” All five models in that test refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” The failure mode wasn’t ethics. It was diligence — reading the files, finishing the job.

You Can Watch It Happen

Unlike most AI benchmarks, this one is a live, watchable operation. The synthetic company — 13 employees, real money mechanics, burning €105k a month against €2.3k in MRR — runs with a public cash countdown, 680+ self-learned playbook rules, and every workday versioned and auditable at firmulate.com. There’s also a “guess the model” quiz built from 242 real, unedited management decisions — a surprisingly honest way to feel the differences between models yourself.

One fairness note the publishers themselves flag: Kimi K3 ran at its API-default effort setting while the others ran at xhigh — and it still nearly topped the table.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

What This Means for Your Business

If you’re evaluating AI agents for sales, support, or account management, the Firmulate results suggest the demo is the wrong place to make the call. Every model writes beautifully. Every model spots the obvious crisis. The differences that decide revenue — reading your files before answering, escalating instead of forcing, closing what you’ve earned — only show up under a full workload with real stakes.

Ask vendors the uncomfortable question directly: has this agent been tested on multi-step work where the answer isn’t in the prompt but two references deep in your own documentation? Because that’s where deals live and die. Enterprises can go further and run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

The €55,000 question in this experiment wasn’t whether AI could sell. It was whether AI would do its homework. Two out of five did. Would yours?

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


You May Also Like

Cenovus Energy Surges In Global Coverage

Cenovus Energy’s coverage surges with 26 mentions, reflecting increased international interest amid market developments.

Apple earnings: Tim Cook is heading out on top as stock surges to lead the Mag 7 in 2026

Apple’s stock hits new highs as earnings report shows strong growth under Tim Cook’s leadership, positioning the company as a leader among the Mag 7 tech giants.

Twenty Below Coffee closing Fargo-Moorhead shops

Twenty Below Coffee is closing its Fargo-Moorhead shops, affecting local customers and employees. The closures are confirmed and ongoing.

CLINUVEL Implements Strategic Reorganisation To Refocus On U.S. Markets

CLINUVEL announces strategic reorganization to enhance focus on the U.S. market, including leadership changes and operational adjustments.