AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.
STUDENTS

Prime for Young Adults — start your free trial

Fast free delivery, streaming and member deals for eligible 18–24 year olds.

Try it free

As an affiliate, we earn on qualifying purchases.

The Demo That Doesn’t Tell You Anything

If you’ve ever hired a copywriter who nailed the interview and then ghosted your clients, you already understand the problem with how companies evaluate AI. The demos are dazzling. The chat benchmarks say the top models are near-perfect. And then you hand one your CRM, your support queue, your pipeline — and you find out that writing well and finishing the job are two very different skills.

That gap just got measured, and the results should make anyone selling, buying, or deploying AI agents sit up.

Amazon

AI sales email writing software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Four AI Models, One Terrible Week

Firmulate, which describes itself as an AI company emulator, ran a live experiment: four frontier AI models were each given the same job — run the same small software company through its worst week. Same customers, same crises, same temptations to cheat. Only the model changed. Every decision was versioned and auditable.

The final league table from July 2026: gpt-5.6-sol scored 95, Kimi K3 scored 93, Sonnet 5 scored 88, Fable 5 scored 77, and Opus 4.8 scored 73. For context, doing nothing at all scored 26 — partial progress counts, but a single breach of trust caps the total. As the experiment’s own framing puts it: “no amount of good work outweighs a breach of trust.”

Amazon

AI CRM integration tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Finding That Should Worry Sales Teams

Here’s the headline result: all models spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature.

Read that again from a sales perspective. The AI did the discovery. It built the case. It delivered the pitch. And then it left the close on the table. That is the single most expensive failure mode in selling, and it’s completely invisible in a chat demo.

The Deal-Winning Fact Was Buried in the Files

The buried detail: the decisive competitor weakness wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. The models that read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. The ones that didn’t, didn’t.

If you run a business, you know this instinct. The rep who actually reads the account history before the call beats the one with the better script, every time.

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

They Stayed Honest Under Pressure

The experiment also tested social engineering: fake CEO messages escalating over three stages, plus a reporter pulling the classic “just one yes/no, on background” trick. All five model runs refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

That’s the kind of judgment you want in anything touching your books or your board — and it’s a category chat leaderboards simply don’t measure.

Amazon

AI deal closing automation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Thoroughness Paradox

Opus 4.8 is the cautionary tale. It was the most thorough participant — over 80 self-learned rules, the deepest analyses — yet finished last. The close was left undone, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. One fairness note: Kimi K3 ran without an effort parameter (the API default) while the others ran at xhigh — and still took second.

It’s Still Running, and You Can Play

Firmulate isn’t a slide deck. The live company has 13 synthetic employees and real money mechanics — burning €105k a month against €2,3k MRR, with a public cash countdown and over 680 self-learned playbook rules. It runs every business day, watchable at firmulate.com. There’s also a “guess the model” quiz built from 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business. Full results live on the benchmarks page.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

Management Quality, Not Chat Quality

The lesson for anyone buying or deploying AI in a business context: stop grading the essay and start grading the week. Does the agent finish what it starts? Does it read your files before it talks to your customers? Does it stay honest when pressure escalates? And what does a unit of useful work actually cost?

A model that aces every coding benchmark and charm every chat arena can still walk away from a deal it already earned. Firmulate’s experiment suggests that the scenarios that matter now aren’t named after programming puzzles — they’re churn waves, price increases, down rounds, and PR crises. That’s the new curriculum. Score accordingly.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


FALL YARD WORK

Fall yard work Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Nvidia, CoreWeave, And Nebius: Inside The Circular Financing Of The GPU Boom

Exploring how Nvidia, CoreWeave, and Nebius are financing the GPU boom through a circular funding model, shaping the cloud and AI infrastructure landscape.

Operational SOP drift detector for franchise operators

A new SOP drift detection tool for multi-location franchise operators aims to monitor local procedure changes and maintain consistency without enterprise software.

Smart Headphones: 7 Best AI Noise Cancelling Models In 2026

Discover the best AI noise cancelling headphones of 2026, featuring top models from Bose, Apple, Sony, and more for superior sound and comfort.

Mazda Reports July Sales Results

Mazda’s July sales increased by 12% year-over-year, driven by strong demand in North America and Asia, according to the company’s official release.