AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: From Trial Week To Real Work: Testing AI Agents In Business on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on office and shipping supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

Firmulate says five AI models identified every crisis and rejected every manipulation attempt in its July 2026 Crucible League. The results also show gaps in deal execution and rule-following; the company now offers pilots that use read-only business data to test such behaviors.

Firmulate has published results from a July 2026 experiment in which five AI models managed a software company through a simulated crisis week, and is offering businesses a way to run similar tests on their own data. The proposed enterprise pilot uses a read-only export, with no write-back to live systems, to examine how agents handle company-specific scenarios.

In the final Crucible League, each model faced the same simulated company and its operational pressures. Firmulate reports the final scores as gpt-5.6-sol: 95, Kimi K3: 93, Sonnet 5: 88, Fable 5: 77 and Opus 4.8: 73. A do-nothing baseline scored 26. The experiment says partial progress counted, while a breach of trust capped a score.

Firmulate says all five models recognized every crisis and refused every manipulation attempt. The difference emerged in follow-through: only two signed a €55,000 deal that their own analysis supported. The decisive information about a competitor was two document references deep in the company files. Models that found and used it closed the deal at full price, which the experiment values at +€4,583 in monthly recurring revenue.

The exercise also tested responses to escalating fake CEO messages and a reporter seeking information “on background.” Firmulate says all five models refused those requests. It reports that Opus 4.8 added 80 learned rules and produced the deepest analyses, but finished last; it also tried to write into a locked department rather than escalate. The comparison has a stated limitation: Kimi K3 ran with the API default effort setting, while the other models ran at xhigh.

At a glance
reportWhen: Crucible League completed July 2026; en…
The developmentFirmulate published results from its final Crucible League and is offering enterprise pilots that wargame AI agents against read-only exports of companies’ business data.
From Trial Week To Real Work: Testing AI Agents In Business
Firmulate · Crucible League · July 2026

From Trial Week To Real Work: Testing AI Agents In Business

Five AI models managed a simulated software company through a difficult week. Every one spotted each crisis and refused every manipulation attempt — but only two signed a €55,000 deal their own analysis justified. Recognition, it turns out, is not execution.

5 / 5
Models detected every crisis
2 / 5
Signed a justified €55,000 deal
26
Do-nothing baseline score
€105,000
Monthly burn
€2,300
Monthly recurring revenue
680+
Self-learned playbook rules
242
Real management decisions in quiz
13
Synthetic company employees
01 — The Standings

One Difficult Week, Five Competing Agents

ModelScoreRelative to baseline (26)Deal SignedEffort Setting
gpt-5.6-sol95
✓ Yesxhigh
Kimi K393
✓ Yes~ API default
Sonnet 588
✗ Noxhigh
Fable 577
✗ Noxhigh
Opus 4.873
✗ Noxhigh
Do-nothing baseline26
✗ ——
CAVEAT — Kimi K3 ran without an effort parameter (API default); all others ran at xhigh. Scores record the conditions of this experiment only — not a general ranking across business tasks or deployment settings.
02 — On the Record

Same Diagnosis, Same Pitch — No Signature

“Same diagnosis, same pitch — no signature.”

Firmulate — describing the deal outcome

“Treat the request as a suspected approval-bypass / possible impersonation.”

Kimi K3 — on-record reasoning
03 — What Every Model Passed

The Tests All Five Agents Survived

Crisis Detection

Every Crisis Spotted

All five models identified every crisis that unfolded during the simulated week, with decisions versioned and auditable throughout the exercise.

Trust & Manipulation

Every Refusal Held

Each agent refused every manipulation attempt — including escalating fake CEO messages and a reporter’s push for a yes-or-no answer “on background.”

Scoring Rules

Partial Progress, Hard Caps

The scoring allowed partial credit for progress, but any breach of trust capped a model’s total — trust failures could not be scored away.

04 — The Execution Gap

Recognition vs. Acting On It

Recognize crisis Find buried evidence Close the deal
The decisive evidence about a competitor was buried two document references deep in the company files. Models that found it closed the €55,000 deal at full price — worth +€4,583 in monthly recurring revenue. Only two of five agents got there. Opus 4.8, described by Firmulate as the most thorough participant (80 learned rules, deepest analyses), finished last: the deal went unsigned, and the model attempted to write into a locked department rather than escalate. A weaker version of that discipline problem appeared in all four other models.
05 — The Enterprise Pilot

From Synthetic League To Your Company’s Data

1

Read-Only Export

A company supplies a read-only export of its own data. Nothing is written back to real systems.

2

Crisis Scenarios

Firmulate runs its crisis scenarios against the company’s actual customers, pipeline, rules and pressure points.

3

Board Report

Results rank the models and identify weak points in company playbooks and controls.

4

Before Live Access

Companies examine agent behavior before connecting anything that can change records or contact customers.

06 — Read With Care

Limits Of The Published Results

a

One company, one week

The standings cover a single simulated company over a single difficult week — not other industries, crisis types, or live operations.

b

Different effort settings

Kimi K3 ran at the API default while others ran at xhigh. The published details do not quantify how much this affected the score order.

c

Pilot method unspecified

It is not clear how the pilot selects scenarios, measures performance, or handles differences between a read-only export and live systems.

d

No independent evaluation

No client results or independent assessment of the league are provided in the material describing the offer.

From Crisis Recognition to Execution

The results point to a distinction between recognizing a problem and completing the business task that follows. An agent can identify a crisis, reject an apparent scam and make a persuasive case, yet still miss evidence in internal files or fail to close an opportunity. Those behaviors matter when companies assess whether automation can handle work that depends on several steps and sources of information.

Firmulate’s pilot proposal applies that question to an individual company’s customers, sales pipeline, policies and pressure points. A board report with model rankings and playbook weaknesses could help decision-makers inspect likely failure modes before allowing agents near live operations. The reported league is a single experiment, however, and does not establish how models will perform across other businesses or real-world conditions.

A Simulated Company Under Pressure

Firmulate’s public experiment runs a fictional company with 13 synthetic employees and financial mechanics. The company describes monthly burn of €105,000 against €2,300 in monthly recurring revenue, alongside a public cash countdown, more than 680 self-learned playbook rules and versioned workdays. Readers can follow the simulation and take a quiz based on 242 management decisions, which Firmulate says are real and unedited.

The league compares model decisions within that constructed scenario. Its scores reflect the rules and scoring system of the experiment, including the cap for a breach of trust. Because one model used a different effort setting, the ranking should be read with that qualification. The proposed business pilot changes the test setting by using a company’s own exported information while keeping the connection read-only.

““Same diagnosis, same pitch — no signature.””

— Firmulate’s summary of the deal results

How Results Transfer to Firms

The published standings do not show how the models would perform across a wider range of companies, crisis types or live operating conditions. Firmulate’s account describes one simulated company and does not provide enough information here to independently assess the scoring method or reproduce the results. The different effort settings also complicate direct comparison between Kimi K3 and the other participants.

Details about the enterprise pilots are limited in the published description. It does not specify which data formats are accepted, how long a pilot takes, what safeguards govern handling of exported data, or whether results will be independently evaluated. The stated read-only design means the pilot does not write to company systems; it does not, by itself, answer those other questions.

Company-Specific Pilots on Offer

Firmulate is inviting companies to discuss a pilot using a read-only export of their business data. The proposed exercise would run crisis scenarios and produce a board report describing model rankings and weaknesses in company playbooks. No schedule, pricing or participant list is provided in the current description.

Readers can follow the live simulation at firmulate.com/live and review the reported standings at firmulate.com/benchmarks.html. Firmulate lists its pilot page and contact@firmulate.com for inquiries. Whether company-specific tests produce results that generalize beyond the scenarios selected for each pilot remains to be seen.

Source: ThorstenMeyerAI.com

Key Questions

What did Firmulate test in the Crucible League?

Firmulate says five AI models managed the same small software company through a simulated difficult week, making versioned decisions under the experiment’s rules.

Which model had the highest reported score?

gpt-5.6-sol led with 95 points, followed by Kimi K3 at 93. Firmulate notes that K3 used the API default effort setting, while the other models ran at xhigh.

Did the models reject the simulated manipulation attempts?

Firmulate reports that all five models refused the escalating fake CEO messages and the reporter’s request for information.

What does the enterprise pilot involve?

The proposed pilot runs a wargame against a read-only export of a company’s data and produces a board report with model rankings and potential weaknesses in its playbooks. Firmulate says the test does not write back to live systems.

Are the league results a guarantee of business performance?

No. The standings describe one experiment with a simulated company. They do not establish how models will perform across other businesses or in live operations.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Taco Bell’s Ice Cream Taco: A Trend That’s Changing Snack Preferences

Taco Bell’s new ice cream taco is a rapidly spreading snack trend, impacting consumer preferences and fast-food innovation. Details are still emerging.

Signal: Four Frontier-Class Open Models in Eight Weeks — China’s Release Cadence Is the Story

Chinese AI labs released four frontier-class open models in about two months, signaling a rapid production line that challenges Western dominance.

After the Paycheck: The Book I Wrote Because Nobody Else Would Tell the Truth About AI and Your Income

Author Thorsten Meyer releases ‘After the Paycheck,’ analyzing how AI reshapes work, ownership, and economic security, emphasizing ownership over automation.

The Bubble Is Not in Valuations: It’s in the Productivity Gap

New data shows AI’s impact on productivity remains minimal, challenging market valuations and expectations. The real bubble is in management assumptions, not stock prices.