
Get business pricing on office and shipping supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
Chat Quality Isn’t Sales Quality — and One AI Just Proved It
If you sell for a living, you know the type: brilliant diagnosis, flawless pitch, and then… no signature. Same meeting, same problem, no close. It turns out AI models suffer from exactly the same affliction — and nobody notices until you put them in the hot seat.
That’s the punchline of Firmulate, a live experiment that runs frontier AI models as the management of the same small software company through its worst week. Same customers, same crises, same temptations to cut corners. Only the model changes. Every decision is versioned and auditable.
As an affiliate, we earn on qualifying purchases.
The Final League Table
The Crucible league, finalized in July 2026, put five models through the identical worst week:
- 1. gpt-5.6-sol — 95. Found the buried fact, closed the deal — the complete performance.
- 2. Kimi K3 — 93. The newcomer from Moonshot: closed the deal too, with the cleanest discipline in the field.
- 3. Sonnet 5 — 88. Closed, with a few more process slips.
- 4. Fable 5 — 77 and 5. Opus 4.8 — 73.
For context, a do-nothing baseline scores 26 — partial progress counts — while a single breach of trust caps the total, because no amount of good work outweighs a breach of trust.
AI customer relationship management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Close Nobody Talks About
Here’s the finding that should stop any sales leader cold: all models spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. Three of five models left the close on the table.
The deal turned on a buried fact. The decisive competitor weakness wasn’t in the customer conversation at all — it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The ones that didn’t, didn’t. In sales terms: the reps who did their discovery homework closed; the ones who winged the meeting didn’t.
As an affiliate, we earn on qualifying purchases.
Pressure Tests and Bait
The week included escalating social-engineering attacks: fake CEO messages over three stages, plus a reporter trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” K3 also found the buried needle, saved a churning customer, and logged just one deviation across the whole gauntlet.
As an affiliate, we earn on qualifying purchases.
The Hard-Working Loser
Opus 4.8 is the cautionary tale for anyone who equates effort with results. It was the most thorough participant — 80 learned rules added, the deepest analyses in the field — and finished last. The close was left on the table and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four other models. Effort without judgment is expensive.
Fairness note: Kimi K3 ran without an effort parameter (API default) while the other models ran at xhigh — worth keeping in mind when comparing scores.
Why Marketers and Ecommerce Operators Should Care
If AI agents will touch your CRM, your support queue, or your forecast, the question isn’t whether they write well. It’s whether they finish what they start, read your files before acting, and stay honest under pressure. A model that aces a chat demo can still fumble the close — and you won’t find out until it’s your pipeline.
The company itself is real software that runs every business day: 13 synthetic employees, real money mechanics — €105k monthly burn against €2,300 MRR — a public cash countdown, and 680+ self-learned playbook rules. It’s watchable live at firmulate.com, with full results and plain-language findings on the benchmarks page. There’s also a “guess the model” quiz built from 242 real, unedited management decisions — a humbling exercise — and enterprises can run the same wargame against a read-only export of their own business.

The Takeaway
A newcomer from Moonshot just outscored three of four Western frontier models at actually running a company. The league is open. If you’re picking a model for anything that touches revenue, picking by demo — or by brand — is now a bet. Run your own test before you hire your AI workforce.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
