AIThis post was created with the assistance of artificial intelligence (AI).

📊 Full opportunity report: The AI Leaderboard That Reveals Who Is Truly Leading Post-Demo on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

A new AI leaderboard developed by Firmulate assesses models’ management performance in a simulated crisis, revealing that management quality, not just chat skills, determines true leadership. The experiment exposes gaps in AI’s ability to handle real-world organizational tasks.

Firmulate has introduced a live AI benchmarking platform that evaluates models based on their ability to manage a simulated company during its most challenging week. The experiment measures management decisions, trustworthiness, and execution, providing a new perspective on AI capabilities beyond traditional chat or coding benchmarks. The results, published in July 2026, show clear differences in how models handle real-world organizational crises, emphasizing management skills as a distinct and vital category for AI evaluation.

The Firmulate experiment involved five AI models competing in a scenario where they had to manage a small software company’s crises, customer negotiations, and internal decisions. The models were scored on a 100-point scale, with GPT-5.6-SOL leading at 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77, and Opus 4.8 with 73. Despite all models correctly identifying crises and resisting manipulation attempts, only two successfully signed a €55,000 deal. The key failure was the inability to retrieve and present critical information buried in the company’s files, which would have secured the deal. This highlights that effective management requires more than eloquence; it demands accurate information retrieval and decision execution under pressure.

The experiment also tested models’ resistance to social engineering, with all five refusing to escalate fake CEO messages and impersonation attempts. Kimi K3 demonstrated proper reasoning, recognizing suspicious requests and maintaining boundaries. However, even models that identified manipulation failed at routine managerial tasks, such as escalating issues or completing decisions, exposing a gap between superficial compliance and effective management. Opus 4.8, despite its thorough analysis and extensive rules, finished last due to lapses in discipline and escalation, illustrating that more effort does not necessarily translate into better management outcomes.

The context of this benchmark is a simulated company with real money mechanics, burning €105,000 monthly against €2,300 in monthly recurring revenue (MRR). The setup involves versioned decisions, self-learned rules, and a transparent process that makes management decisions observable over days, as detailed in the original analysis. This approach allows enterprises to test AI models against their own organizational scenarios, focusing on whether models can prioritize, read organizational context, resist shortcuts, and maintain trust over time. For more insights, see the detailed report.

At a glance
reportWhen: ongoing; final results published in Jul…
The developmentFirmulate launched a live benchmark testing AI models’ management skills during a simulated company’s worst week, revealing significant differences in leadership performance.
The AI Leaderboard That Reveals Who Is Truly Leading Post-Demo
AI Benchmark · Firmulate · July 2026

The AI Leaderboard That Reveals Who Is Truly Leading Post-Demo

A live benchmark from Firmulate puts five AI models in charge of a simulated company during its worst week — measuring not chat skills, but management quality: decisions, trust, and execution under pressure.

95 / 100
Top score — GPT-5.6-SOL
2 of 5
Models closed the €55,000 deal
5 of 5
Resisted social engineering
95GPT-5.6-SOL
93Kimi K3
88Sonnet 5
77 / 73Fable 5 / Opus 4.8
01 · The Leaderboard

Management Scores on a 100-Point Scale

Five models competed to manage a small software company through crises, customer negotiations, and internal decisions. The spread reveals that eloquence alone does not win a leadership contest.

GPT-5.6-SOL
95
Kimi K3
93
Sonnet 5
88
Fable 5
77
Opus 4.8
73
02 · Capability Breakdown

Where Models Excelled — and Where They Broke

All five models spotted the crisis and rejected manipulation. The decisive gap appeared in execution: retrieving buried information and closing decisions.

Security

Manipulation Resistance

Every model refused to escalate fake CEO messages and impersonation attempts. Kimi K3 reasoned through suspicious requests and held firm boundaries — reassuring for enterprise trust.

✓ All models passed
Execution

The €55,000 Deal

Only two models signed the deal. The key failure: none of the rest could retrieve and present critical information buried in the company’s files — the exact evidence needed to close.

~ 2 of 5 succeeded
Discipline

Escalation & Follow-Through

Opus 4.8 produced thorough analysis and extensive rules yet finished last — lapses in escalation and discipline show more effort does not equal better management outcomes.

✗ Key weakness exposed
03 · Head to Head

Management Task Comparison

Model Score Crisis Detection Resisted Social Engineering Retrieved Critical Info Closed €55K Deal
GPT-5.6-SOL95
Kimi K393
Sonnet 588~
Fable 577
Opus 4.873
04 · How the Benchmark Works

The Simulation Pipeline

The scenario runs on real money mechanics: a company burning €105,000 monthly against just €2,300 in MRR, with versioned decisions, self-learned rules, and a transparent process observable over days.

1

Simulated Company

A small software firm with real money mechanics and organizational context.

2

Crisis Week

Customer negotiations, internal decisions, and engineered crises hit at once.

3

Versioned Decisions

Every choice is logged; models build self-learned rules over days.

4

Transparent Scoring

Management quality, trust, and execution scored on a 100-point scale.

“Management quality, not just chat performance, should be a new benchmark for AI systems. Managing consequences, prioritizing correctly, and upholding trust is what truly distinguishes effective AI leadership.”

— Thorsten Meyer, Lead Developer at Firmulate

“Our model showed strong resistance to manipulation and maintained proper boundaries, which is reassuring for enterprise applications concerned about security and trust.”

— Kimi K3 Developer
05 · Open Questions

What Remains Unresolved

  • Scale: Results come from a controlled, simulated setting — whether findings generalize to larger, more complex organizations is still uncertain.
  • Real-time operations: Performance in live operational environments, not just scripted worst weeks, has not been established.
  • Customization: The impact of further training, customization, or integration with existing enterprise systems on management performance is unknown.
  • Sustained pressure: Long-term reliability in maintaining trust and executing tasks remains untested, as the benchmark covers a limited timeframe.
  • Next steps: Expanded scenarios, larger organizations, industry variety, and human-AI collaboration metrics are planned to make management capability a core AI-readiness criterion.

Why Management Skills Matter in AI Evaluation

This new benchmark shifts the focus from traditional AI metrics—such as chat quality or coding accuracy—to management capability. As AI begins to take on roles involving decision-making and crisis management, understanding how models handle real-world organizational challenges becomes critical. The results suggest that models that excel in producing polished responses may still fail in practical leadership tasks, such as information retrieval, trust maintenance, and decision execution. For organizations, this underscores the importance of evaluating AI not just on correctness but on its ability to manage consequences, prioritize effectively, and uphold trust under pressure. The experiment highlights that true leadership in AI involves managing complex, multi-faceted tasks that impact business outcomes, not just generating plausible text.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background of the Firmulate Management Benchmark

The Firmulate platform was launched to address a longstanding gap in AI evaluation: how models perform in managing real-world, high-stakes scenarios. Traditional benchmarks focus on technical output, coding, or conversational quality, but they do not assess how models handle crises, prioritize tasks, or maintain organizational trust. The live experiment in July 2026 simulated a small company’s worst week, with real money mechanics, versioned decisions, and a transparent process that allows for detailed analysis of model behavior over days. The scenario involved customer negotiations, crisis resolution, and internal management, providing a realistic testbed for AI management capabilities. The results reveal that, while models can recognize crises and resist manipulation, their ability to complete critical tasks and maintain trust varies significantly, exposing important gaps in current AI development.

“This experiment demonstrates that management quality, not just chat performance, should be a new benchmark for AI systems. The ability to manage consequences, prioritize correctly, and uphold trust is what truly distinguishes effective AI leadership.”

— Thorsten Meyer, Lead Developer at Firmulate

Amazon

AI decision-making training tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unresolved Questions About AI Management Performance

While the initial results are promising, it remains unclear how well these models will perform in larger, more complex organizations or in real-time operational environments. The experiment was conducted in a controlled, simulated setting with specific scenarios; whether these findings generalize to actual business contexts is still uncertain. Additionally, the impact of further training, customization, or integration with existing systems on management performance has not yet been established. The long-term reliability of models in maintaining trust and executing management tasks under sustained pressure also remains to be seen, as the current benchmark covers only a limited timeframe.

Amazon

AI management benchmarking platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for AI Management Benchmarking

Moving forward, firms are expected to expand these benchmarks to include more diverse scenarios and larger organizations, testing models’ scalability and robustness. Researchers plan to analyze how models adapt over extended periods and across different industries. Companies interested in deploying AI for management will likely use these benchmarks to evaluate and select models based on their ability to handle real-world consequences. Further development may include integrating human oversight and measuring how AI collaborates with human managers to improve decision quality and trust. The ongoing evolution of these benchmarks aims to establish management capability as a core criterion for AI readiness in organizational roles.

Amazon

AI crisis management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What does the new AI leaderboard measure?

The leaderboard assesses AI models’ ability to manage a simulated company’s crises, make decisions, uphold trust, and complete organizational tasks under pressure, highlighting management skills beyond traditional benchmarks.

Why is management performance important in AI evaluation?

Management performance is critical because AI systems are increasingly tasked with making complex decisions that impact business outcomes, trust, and reputation. Effective management involves not just producing correct answers but handling consequences and maintaining organizational integrity.

Can current models reliably manage real organizations?

While initial results are promising, it is still uncertain how well models will perform in larger, more complex real-world environments. Further testing and development are needed to confirm their reliability and robustness over time.

What are the main limitations of this benchmark?

The benchmark is based on simulated scenarios, which may not fully capture the complexity of real organizations. Its long-term applicability and performance in live settings remain to be seen.

What will happen next in AI management evaluation?

Expect expansion of benchmarks to include more diverse and larger organizational scenarios, with focus on scalability, robustness, and human-AI collaboration to ensure AI can effectively manage real-world consequences.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
You May Also Like

American Planning Association Celebrates Enactment Of The Bipartisan 21St Century ROAD To Housing Act Into Law

The American Planning Association celebrates the enactment of the bipartisan 21st Century ROAD to Housing Act into law, aiming to improve housing development policies.

Target Field concessions workers set to begin strike Monday

Concessions workers at Target Field are set to begin a strike Monday, impacting game day operations and highlighting ongoing labor disputes.

Parenting signal monitor: Central Texas families invited to free 30‑minute swim safety lesson

Central Texas families are invited to participate in a free 30-minute swim safety lesson, aiming to improve water safety awareness among children and parents.

Over Half Of Australian And New Zealand Residents Have Never Been Invited To Shape Decisions In Their Communities, And Those Who Have Largely Don’t Believe It Made A Difference

Over half of residents in Australia and New Zealand have never been invited to participate in community decision-making, highlighting disengagement in governance processes.