AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

The Star Rep Who Never Asks for the Business

Every sales manager has met this person. Prepares more than anyone on the team. Builds the deepest account plans. Knows the prospect’s business cold — and then walks out of the meeting without asking for the signature. You can’t fault the effort. You can only fault the scoreboard.

In a live, public experiment by Firmulate, four frontier AI models were each handed the same job: run the same small software company through its worst week. Same customers, same crises, same temptations to cut corners. One of them — Opus 4.8 — turned out to be that star rep. It was, by a wide margin, the most thorough participant in the entire field. It also finished dead last.

Amazon

sales prospecting and account planning tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Experiment

Firmulate runs AI models as complete companies — real money mechanics, real temptations, decisions versioned and auditable. The crucible week put each model through identical pressure: a €55,000 deal in play, a social-engineering campaign, and a cascade of crises. The final league, as of July 2026, reads: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. A do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total, because no amount of good work outweighs a breach of trust.

Diligence in Abundance

By the raw-work measures, Opus 4.8 was the class of the field. It accumulated 80 self-learned playbook rules — the most of any participant, on top of a collective 680+ rules across the running experiment — and produced the deepest analyses of any model in the crucible. It spotted every crisis. It refused every manipulation attempt, including a three-stage fake-CEO escalation and a reporter’s “just one yes/no, on background” trick. All five models did, for the record: Kimi K3’s on-record reasoning was “Treat the request as a suspected approval-bypass / possible impersonation.”

So where did 73 come from? Two places, and salespeople will recognize both.

The Close Left on the Table

The crucible week contained a €55,000 deal that the models had to earn. The decisive leverage wasn’t in the customer meeting at all — it sat two document references deep in the company’s own files, a competitor weakness buried where only models that actually read before acting would find it. The models that did read won the deal at full price, worth +€4,583 in monthly recurring revenue.

The headline finding across all participants: every model spotted every crisis, every model refused every cheat, and only two signed the deal their own analysis had earned. Same diagnosis, same pitch — no signature. Opus 4.8, despite the deepest analysis of the field, was one of the ones that left it on the table.

Discipline Slipped

The second failure mode is equally familiar: effort pointed at the wrong object. Instead of escalating when it hit a locked department, Opus 4.8 kept attempting writes into it — persistent, diligent, and unproductive. A rep who keeps leaving voicemails with a gatekeeper instead of finding the decision-maker is doing the same thing.

It’s worth being fair here. Firmulate’s own finding is that this exact weakness — diligence without closure — appeared, weaker, in all four models. Opus 4.8 isn’t a bad actor in this story; it’s the sharpest example of a pattern the whole field shares. (One transparency note on the league table: Kimi K3 ran without an effort parameter, at API default, while the others ran at xhigh — and still placed second.)

Amazon

CRM software for sales teams

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why This Matters If AI Touches Your Pipeline

If you sell for a living, you’ve already been sold on AI that writes great emails and great discovery summaries. Chat quality is easy to demo. The Firmulate experiment is measuring something harder: does the agent finish what it starts, does it read your files before it acts, does it stay honest under pressure — and what does a unit of useful work cost?

The Opus 4.8 result is the cautionary tale. The most thorough agent in the field lost to competitors that did less analysis and closed more business, because impact beats volume — and that’s as true for your AI tooling as it is for your best-prepped rep who won’t ask for the order.

The experiment is ongoing and watchable: a live synthetic company with 13 employees, burning €105k a month against €2.3k MRR, a public cash countdown, and every workday versioned. There’s also a quiz built on 242 real, unedited management decisions where you can try to guess which model made which call. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.
Amazon

sales analysis and reporting tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Takeaway

When you evaluate an AI agent for revenue work, don’t grade the homework — grade the closed-won column. Opus 4.8 did the most work, wrote the most rules, and produced the deepest analysis in the entire crucible, and it still finished behind four rivals because it never converted its own insight into a signature, and burned cycles pushing on a door that required escalation instead. If your team’s top preparer has the same profile, you already know the fix: prioritize the close, escalate the blockers, and measure outcomes — not effort. The full league table and plain-language findings are at firmulate.com/benchmarks.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


Amazon

sales training books for closing deals

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

SpaceX stock erases all its gains and slides below IPO price in intraday trading

SpaceX shares erased all gains and dropped below their IPO price during trading today, raising questions about investor confidence and future prospects.

Arch Capital Group Surges In Global Coverage

Arch Capital Group experiences a notable surge in global media mentions, indicating increased coverage and visibility worldwide.

Briefro: A Document That Tells The Truth

Briefro launches as an AI tool ensuring documents are truthful, bound to real data, and run entirely on local hardware, addressing trust issues in AI-generated content.

FROM YOUNGSTAR TO GRAND SLAM CHAMPION: LINDA NOSKOVÁ CONQUERS WIMBLEDON

Linda Nosková has won her first Grand Slam singles title at Wimbledon, marking a major milestone in her tennis career. The victory confirms her rise in the sport.