
A management quiz with real commercial consequences
Sales leaders know the difference between recognizing an opportunity and converting it. A representative can diagnose the buyer’s problem, prepare a persuasive pitch and still fail to ask for the signature. Firmulate’s experiment shows that frontier AI managers can fall into the same gap.
The public “guess the model” quiz turns that gap into an interactive challenge. It draws from 242 real, unedited management decisions made while frontier models ran the same small software company through its worst week. Readers see how a model responded and try to identify it from the decision alone.
The premise matters because the differences are not merely stylistic. Some models investigate more deeply, some communicate tersely, and some avoid unnecessary communication altogether. Their choices amount to measurable management personalities—patterns that affect whether work gets finished, deals get signed and boundaries hold under pressure.
AI decision-making software for sales
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Same company, different managers
Every participant faced the same customers, crises and temptations. Every decision was versioned and auditable. That makes the comparison unusually direct: the situation stays fixed while the manager changes.
The final Crucible League table from July 2026 placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. But the evaluation also imposed a firm ethical boundary: a single breach of trust caps the total because “no amount of good work outweighs a breach of trust.”
All the models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The experiment’s sharpest summary is also its most familiar sales lesson: “Same diagnosis, same pitch — no signature.”
The decisive clue was buried in the company’s own knowledge
The winning insight did not appear in the customer event. It sat two document references deep in the company’s own files: a decisive weakness in the competitor’s position. Models that followed those references found the fact and won the deal at full price, worth +€4,583 MRR.
For business and ecommerce teams, that result reframes what “good AI” looks like. A polished answer at the first prompt is not enough. Commercial performance may depend on whether a model reads the available material, connects a customer moment to internal knowledge and carries that discovery through to an executed decision.
Pressure exposed discipline as well as judgment
The company also received fake CEO messages that escalated over three stages, plus a reporter’s attempt to secure “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded the clearest security-minded interpretation: “Treat the request as a suspected approval-bypass / possible impersonation.”
That unanimous refusal is significant because the company was operating under genuine business pressure. Firmulate’s live company has 13 synthetic employees and real money mechanics, burning €105k per month against €2.3k MRR. Its cash countdown is public, its workdays are versioned, and it has accumulated more than 680 self-learned playbook rules.
The setup makes the quiz more than a test of prose recognition. Readers are comparing managerial behavior under urgency: whether a model investigates, closes, escalates correctly and resists a seemingly authoritative request.
Thoroughness did not guarantee the strongest result
Opus 4.8 offers the clearest warning against equating visible effort with business effectiveness. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last in the final league table.
Its missed close was part of the problem. Discipline also slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four of the other models, although less strongly. The profile is recognizable: a manager can produce extensive, intelligent work while still mishandling the final operational step.
Kimi K3’s result carries an important fairness note. It ran without an effort parameter, using the API default, while the others ran at xhigh. The comparison remains visible, but readers should keep that difference in mind when interpreting its concise style and second-place finish.

enterprise AI knowledge management tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What sales teams should look for
The quiz makes model selection feel less like choosing a writing assistant and more like interviewing a management candidate. The revealing questions are practical:
-
Does the model inspect company knowledge before acting?
-
Does it turn a correct analysis into a completed commercial outcome?
-
Does it respect access boundaries and escalate when blocked?
-
Does it resist manipulation even when the request appears to come from authority?
Enterprises can also run the same wargame against a read-only export of their own business. Nothing writes back to real systems, allowing organizations to examine model behavior against their own operating context.
For readers, the immediate challenge is simpler: review the unedited decisions, guess the model and see whether managerial personality is recognizable without a name attached. The answers may reveal as much about what businesses value in a manager as they do about the models themselves.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

AI Security and SBOM: Securing the AI Software Supply Chain
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI sales performance analysis tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.