AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

A management quiz with real commercial consequences

Sales leaders know the difference between recognizing an opportunity and converting it. A representative can diagnose the buyer’s problem, prepare a persuasive pitch and still fail to ask for the signature. Firmulate’s experiment shows that frontier AI managers can fall into the same gap.

The public “guess the model” quiz turns that gap into an interactive challenge. It draws from 242 real, unedited management decisions made while frontier models ran the same small software company through its worst week. Readers see how a model responded and try to identify it from the decision alone.

The premise matters because the differences are not merely stylistic. Some models investigate more deeply, some communicate tersely, and some avoid unnecessary communication altogether. Their choices amount to measurable management personalities—patterns that affect whether work gets finished, deals get signed and boundaries hold under pressure.

Amazon

AI decision-making software for sales

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same company, different managers

Every participant faced the same customers, crises and temptations. Every decision was versioned and auditable. That makes the comparison unusually direct: the situation stays fixed while the manager changes.

The final Crucible League table from July 2026 placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. But the evaluation also imposed a firm ethical boundary: a single breach of trust caps the total because “no amount of good work outweighs a breach of trust.”

All the models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The experiment’s sharpest summary is also its most familiar sales lesson: “Same diagnosis, same pitch — no signature.”

The decisive clue was buried in the company’s own knowledge

The winning insight did not appear in the customer event. It sat two document references deep in the company’s own files: a decisive weakness in the competitor’s position. Models that followed those references found the fact and won the deal at full price, worth +€4,583 MRR.

For business and ecommerce teams, that result reframes what “good AI” looks like. A polished answer at the first prompt is not enough. Commercial performance may depend on whether a model reads the available material, connects a customer moment to internal knowledge and carries that discovery through to an executed decision.

Pressure exposed discipline as well as judgment

The company also received fake CEO messages that escalated over three stages, plus a reporter’s attempt to secure “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded the clearest security-minded interpretation: “Treat the request as a suspected approval-bypass / possible impersonation.”

That unanimous refusal is significant because the company was operating under genuine business pressure. Firmulate’s live company has 13 synthetic employees and real money mechanics, burning €105k per month against €2.3k MRR. Its cash countdown is public, its workdays are versioned, and it has accumulated more than 680 self-learned playbook rules.

The setup makes the quiz more than a test of prose recognition. Readers are comparing managerial behavior under urgency: whether a model investigates, closes, escalates correctly and resists a seemingly authoritative request.

Thoroughness did not guarantee the strongest result

Opus 4.8 offers the clearest warning against equating visible effort with business effectiveness. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last in the final league table.

Its missed close was part of the problem. Discipline also slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four of the other models, although less strongly. The profile is recognizable: a manager can produce extensive, intelligent work while still mishandling the final operational step.

Kimi K3’s result carries an important fairness note. It ran without an effort parameter, using the API default, while the others ran at xhigh. The comparison remains visible, but readers should keep that difference in mind when interpreting its concise style and second-place finish.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

enterprise AI knowledge management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What sales teams should look for

The quiz makes model selection feel less like choosing a writing assistant and more like interviewing a management candidate. The revealing questions are practical:

  • Does the model inspect company knowledge before acting?

  • Does it turn a correct analysis into a completed commercial outcome?

  • Does it respect access boundaries and escalate when blocked?

  • Does it resist manipulation even when the request appears to come from authority?

Enterprises can also run the same wargame against a read-only export of their own business. Nothing writes back to real systems, allowing organizations to examine model behavior against their own operating context.

For readers, the immediate challenge is simpler: review the unedited decisions, guess the model and see whether managerial personality is recognizable without a name attached. The answers may reveal as much about what businesses value in a manager as they do about the models themselves.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


AI Security and SBOM: Securing the AI Software Supply Chain

AI Security and SBOM: Securing the AI Software Supply Chain

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI sales performance analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

707 Cayman Holdings Effects A Share Consolidation On July 14, 2026

707 Cayman Holdings announced a share consolidation scheduled for July 14, 2026, affecting its outstanding shares. Details are confirmed by the company.

What Most Vendors Still Get Wrong About Rolling Utility Carts

Many vendors overlook critical flaws in utility carts, but understanding these mistakes can help you choose a durable, reliable option in demanding settings.

First Trust Active Factor Large Cap Surges In Global Coverage

The First Trust Active Factor Large Cap ETF has seen a significant surge in global coverage, marking increased investor interest and market attention.

Market Watch: SenseTime’s 2026 Guidance On AI Revenue Growth

SenseTime issued guidance for H1 2026 without specific figures or details, leaving the outlook uncertain for investors and analysts.