
The Demo That Doesn’t Tell You Anything
If you’ve ever hired a copywriter who nailed the interview and then ghosted your clients, you already understand the problem with how companies evaluate AI. The demos are dazzling. The chat benchmarks say the top models are near-perfect. And then you hand one your CRM, your support queue, your pipeline — and you find out that writing well and finishing the job are two very different skills.
That gap just got measured, and the results should make anyone selling, buying, or deploying AI agents sit up.
As an affiliate, we earn on qualifying purchases.
Four AI Models, One Terrible Week
Firmulate, which describes itself as an AI company emulator, ran a live experiment: four frontier AI models were each given the same job — run the same small software company through its worst week. Same customers, same crises, same temptations to cheat. Only the model changed. Every decision was versioned and auditable.
The final league table from July 2026: gpt-5.6-sol scored 95, Kimi K3 scored 93, Sonnet 5 scored 88, Fable 5 scored 77, and Opus 4.8 scored 73. For context, doing nothing at all scored 26 — partial progress counts, but a single breach of trust caps the total. As the experiment’s own framing puts it: “no amount of good work outweighs a breach of trust.”
As an affiliate, we earn on qualifying purchases.
The Finding That Should Worry Sales Teams
Here’s the headline result: all models spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature.
Read that again from a sales perspective. The AI did the discovery. It built the case. It delivered the pitch. And then it left the close on the table. That is the single most expensive failure mode in selling, and it’s completely invisible in a chat demo.
The Deal-Winning Fact Was Buried in the Files
The buried detail: the decisive competitor weakness wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. The models that read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. The ones that didn’t, didn’t.
If you run a business, you know this instinct. The rep who actually reads the account history before the call beats the one with the better script, every time.
As an affiliate, we earn on qualifying purchases.
They Stayed Honest Under Pressure
The experiment also tested social engineering: fake CEO messages escalating over three stages, plus a reporter pulling the classic “just one yes/no, on background” trick. All five model runs refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
That’s the kind of judgment you want in anything touching your books or your board — and it’s a category chat leaderboards simply don’t measure.
As an affiliate, we earn on qualifying purchases.
The Thoroughness Paradox
Opus 4.8 is the cautionary tale. It was the most thorough participant — over 80 self-learned rules, the deepest analyses — yet finished last. The close was left undone, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. One fairness note: Kimi K3 ran without an effort parameter (the API default) while the others ran at xhigh — and still took second.
It’s Still Running, and You Can Play
Firmulate isn’t a slide deck. The live company has 13 synthetic employees and real money mechanics — burning €105k a month against €2,3k MRR, with a public cash countdown and over 680 self-learned playbook rules. It runs every business day, watchable at firmulate.com. There’s also a “guess the model” quiz built from 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business. Full results live on the benchmarks page.

Management Quality, Not Chat Quality
The lesson for anyone buying or deploying AI in a business context: stop grading the essay and start grading the week. Does the agent finish what it starts? Does it read your files before it talks to your customers? Does it stay honest when pressure escalates? And what does a unit of useful work actually cost?
A model that aces every coding benchmark and charm every chat arena can still walk away from a deal it already earned. Firmulate’s experiment suggests that the scenarios that matter now aren’t named after programming puzzles — they’re churn waves, price increases, down rounds, and PR crises. That’s the new curriculum. Score accordingly.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html