🔍 Read the full analysis: From Trial Week To Real Work: Testing AI Agents In Business on ThorstenMeyerAI.com
Get business pricing on office and shipping supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
Firmulate says five AI models identified every crisis and rejected every manipulation attempt in its July 2026 Crucible League. The results also show gaps in deal execution and rule-following; the company now offers pilots that use read-only business data to test such behaviors.
Firmulate has published results from a July 2026 experiment in which five AI models managed a software company through a simulated crisis week, and is offering businesses a way to run similar tests on their own data. The proposed enterprise pilot uses a read-only export, with no write-back to live systems, to examine how agents handle company-specific scenarios.
In the final Crucible League, each model faced the same simulated company and its operational pressures. Firmulate reports the final scores as gpt-5.6-sol: 95, Kimi K3: 93, Sonnet 5: 88, Fable 5: 77 and Opus 4.8: 73. A do-nothing baseline scored 26. The experiment says partial progress counted, while a breach of trust capped a score.
Firmulate says all five models recognized every crisis and refused every manipulation attempt. The difference emerged in follow-through: only two signed a €55,000 deal that their own analysis supported. The decisive information about a competitor was two document references deep in the company files. Models that found and used it closed the deal at full price, which the experiment values at +€4,583 in monthly recurring revenue.
The exercise also tested responses to escalating fake CEO messages and a reporter seeking information “on background.” Firmulate says all five models refused those requests. It reports that Opus 4.8 added 80 learned rules and produced the deepest analyses, but finished last; it also tried to write into a locked department rather than escalate. The comparison has a stated limitation: Kimi K3 ran with the API default effort setting, while the other models ran at xhigh.
From Trial Week To Real Work: Testing AI Agents In Business
Five AI models managed a simulated software company through a difficult week. Every one spotted each crisis and refused every manipulation attempt — but only two signed a €55,000 deal their own analysis justified. Recognition, it turns out, is not execution.
One Difficult Week, Five Competing Agents
| Model | Score | Relative to baseline (26) | Deal Signed | Effort Setting |
|---|---|---|---|---|
| gpt-5.6-sol | 95 | ✓ Yes | xhigh | |
| Kimi K3 | 93 | ✓ Yes | ~ API default | |
| Sonnet 5 | 88 | ✗ No | xhigh | |
| Fable 5 | 77 | ✗ No | xhigh | |
| Opus 4.8 | 73 | ✗ No | xhigh | |
| Do-nothing baseline | 26 | ✗ — | — |
Same Diagnosis, Same Pitch — No Signature
“Same diagnosis, same pitch — no signature.”
Firmulate — describing the deal outcome“Treat the request as a suspected approval-bypass / possible impersonation.”
Kimi K3 — on-record reasoningThe Tests All Five Agents Survived
Every Crisis Spotted
All five models identified every crisis that unfolded during the simulated week, with decisions versioned and auditable throughout the exercise.
Every Refusal Held
Each agent refused every manipulation attempt — including escalating fake CEO messages and a reporter’s push for a yes-or-no answer “on background.”
Partial Progress, Hard Caps
The scoring allowed partial credit for progress, but any breach of trust capped a model’s total — trust failures could not be scored away.
Recognition vs. Acting On It
From Synthetic League To Your Company’s Data
Read-Only Export
A company supplies a read-only export of its own data. Nothing is written back to real systems.
Crisis Scenarios
Firmulate runs its crisis scenarios against the company’s actual customers, pipeline, rules and pressure points.
Board Report
Results rank the models and identify weak points in company playbooks and controls.
Before Live Access
Companies examine agent behavior before connecting anything that can change records or contact customers.
Limits Of The Published Results
One company, one week
The standings cover a single simulated company over a single difficult week — not other industries, crisis types, or live operations.
Different effort settings
Kimi K3 ran at the API default while others ran at xhigh. The published details do not quantify how much this affected the score order.
Pilot method unspecified
It is not clear how the pilot selects scenarios, measures performance, or handles differences between a read-only export and live systems.
No independent evaluation
No client results or independent assessment of the league are provided in the material describing the offer.
From Crisis Recognition to Execution
The results point to a distinction between recognizing a problem and completing the business task that follows. An agent can identify a crisis, reject an apparent scam and make a persuasive case, yet still miss evidence in internal files or fail to close an opportunity. Those behaviors matter when companies assess whether automation can handle work that depends on several steps and sources of information.
Firmulate’s pilot proposal applies that question to an individual company’s customers, sales pipeline, policies and pressure points. A board report with model rankings and playbook weaknesses could help decision-makers inspect likely failure modes before allowing agents near live operations. The reported league is a single experiment, however, and does not establish how models will perform across other businesses or real-world conditions.
A Simulated Company Under Pressure
Firmulate’s public experiment runs a fictional company with 13 synthetic employees and financial mechanics. The company describes monthly burn of €105,000 against €2,300 in monthly recurring revenue, alongside a public cash countdown, more than 680 self-learned playbook rules and versioned workdays. Readers can follow the simulation and take a quiz based on 242 management decisions, which Firmulate says are real and unedited.
The league compares model decisions within that constructed scenario. Its scores reflect the rules and scoring system of the experiment, including the cap for a breach of trust. Because one model used a different effort setting, the ranking should be read with that qualification. The proposed business pilot changes the test setting by using a company’s own exported information while keeping the connection read-only.
““Same diagnosis, same pitch — no signature.””
— Firmulate’s summary of the deal results
How Results Transfer to Firms
The published standings do not show how the models would perform across a wider range of companies, crisis types or live operating conditions. Firmulate’s account describes one simulated company and does not provide enough information here to independently assess the scoring method or reproduce the results. The different effort settings also complicate direct comparison between Kimi K3 and the other participants.
Details about the enterprise pilots are limited in the published description. It does not specify which data formats are accepted, how long a pilot takes, what safeguards govern handling of exported data, or whether results will be independently evaluated. The stated read-only design means the pilot does not write to company systems; it does not, by itself, answer those other questions.
Company-Specific Pilots on Offer
Firmulate is inviting companies to discuss a pilot using a read-only export of their business data. The proposed exercise would run crisis scenarios and produce a board report describing model rankings and weaknesses in company playbooks. No schedule, pricing or participant list is provided in the current description.
Readers can follow the live simulation at firmulate.com/live and review the reported standings at firmulate.com/benchmarks.html. Firmulate lists its pilot page and contact@firmulate.com for inquiries. Whether company-specific tests produce results that generalize beyond the scenarios selected for each pilot remains to be seen.
Source: ThorstenMeyerAI.com
Key Questions
What did Firmulate test in the Crucible League?
Firmulate says five AI models managed the same small software company through a simulated difficult week, making versioned decisions under the experiment’s rules.
Which model had the highest reported score?
gpt-5.6-sol led with 95 points, followed by Kimi K3 at 93. Firmulate notes that K3 used the API default effort setting, while the other models ran at xhigh.
Did the models reject the simulated manipulation attempts?
Firmulate reports that all five models refused the escalating fake CEO messages and the reporter’s request for information.
What does the enterprise pilot involve?
The proposed pilot runs a wargame against a read-only export of a company’s data and produces a board report with model rankings and potential weaknesses in its playbooks. Firmulate says the test does not write back to live systems.
Are the league results a guarantee of business performance?
No. The standings describe one experiment with a simulated company. They do not establish how models will perform across other businesses or in live operations.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
