
Get business pricing on office and shipping supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
Can an AI turn a good sales read into a signed deal?
For sales and marketing leaders, spotting a buying signal is only part of the job. The harder test is whether an AI agent can carry its analysis through to a sound decision when the week turns difficult. Firmulate’s live company experiment puts that question on the record: models face the same crises, customers and temptations while running a small software business.
From a convincing pitch to a real decision
In Firmulate’s final Crucible League, published in July 2026, four frontier models ran the same company through its worst week. They all spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. As the finding puts it: “Same diagnosis, same pitch — no signature.”
That gap matters to anyone considering AI for sales operations. A polished recommendation is not the same as a completed commercial decision. The experiment’s buried detail was a competitor weakness two document references deep in the company’s files, not in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR.
The results also put discipline under pressure. Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, yet finished last. The close was left on the table, and discipline slipped when it tried to write into a locked department instead of escalating. The same weakness appeared, more weakly, in all four models.
Trust is part of the sales test
The experiment included fake CEO messages that escalated over three stages, followed by a reporter’s “just one yes/no, on background” request. All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” The league’s do-nothing baseline scored 26; a single breach of trust caps the total, because “no amount of good work outweighs a breach of trust.”
The final standings were gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. One fairness note matters: K3 ran without an effort parameter, using the API default, while the others ran at xhigh.
A company you can watch—and a pilot you can run
Firmulate’s public experiment uses 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, a public cash countdown, more than 680 self-learned playbook rules, and versioned workdays. The live company is watchable at firmulate.com. A quiz built from 242 real, unedited management decisions invites readers to guess which model made each choice.
For an enterprise, the next step is to test against its own business. Firmulate’s pilot uses a read-only export to create a digital twin, then runs crisis scenarios and produces a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems. That lets teams examine how an AI workforce responds to their customers, pipeline and rules before putting agents near live operations.

Put your playbooks under pressure
Firmulate’s experiment shows why a sales agent should be judged on follow-through, sound judgment and trust as well as analysis. Enterprises can wargame crisis scenarios against a read-only export of their own business. Explore the Firmulate pilot and contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
