AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Buying for a business?Offer from Amazon

Get business pricing on office and shipping supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

A new AI leaderboard developed by Firmulate assesses models’ management performance in a simulated crisis, revealing that management quality, not just chat skills, determines true leadership. The experiment exposes gaps in AI’s ability to handle real-world organizational tasks.

Firmulate has introduced a live AI benchmarking platform that evaluates models based on their ability to manage a simulated company during its most challenging week. The experiment measures management decisions, trustworthiness, and execution, providing a new perspective on AI capabilities beyond traditional chat or coding benchmarks. The results, published in July 2026, show clear differences in how models handle real-world organizational crises, emphasizing management skills as a distinct and vital category for AI evaluation.

The Firmulate experiment involved five AI models competing in a scenario where they had to manage a small software company’s crises, customer negotiations, and internal decisions. The models were scored on a 100-point scale, with GPT-5.6-SOL leading at 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77, and Opus 4.8 with 73. Despite all models correctly identifying crises and resisting manipulation attempts, only two successfully signed a €55,000 deal. The key failure was the inability to retrieve and present critical information buried in the company’s files, which would have secured the deal. This highlights that effective management requires more than eloquence; it demands accurate information retrieval and decision execution under pressure.

The experiment also tested models’ resistance to social engineering, with all five refusing to escalate fake CEO messages and impersonation attempts. Kimi K3 demonstrated proper reasoning, recognizing suspicious requests and maintaining boundaries. However, even models that identified manipulation failed at routine managerial tasks, such as escalating issues or completing decisions, exposing a gap between superficial compliance and effective management. Opus 4.8, despite its thorough analysis and extensive rules, finished last due to lapses in discipline and escalation, illustrating that more effort does not necessarily translate into better management outcomes.

The context of this benchmark is a simulated company with real money mechanics, burning €105,000 monthly against €2,300 in monthly recurring revenue (MRR). The setup involves versioned decisions, self-learned rules, and a transparent process that makes management decisions observable over days, as detailed in the original analysis. This approach allows enterprises to test AI models against their own organizational scenarios, focusing on whether models can prioritize, read organizational context, resist shortcuts, and maintain trust over time. For more insights, see the detailed report.

At a glance
reportWhen: ongoing; final results published in Jul…
The developmentFirmulate launched a live benchmark testing AI models’ management skills during a simulated company’s worst week, revealing significant differences in leadership performance.

Why Management Skills Matter in AI Evaluation

This new benchmark shifts the focus from traditional AI metrics—such as chat quality or coding accuracy—to management capability. As AI begins to take on roles involving decision-making and crisis management, understanding how models handle real-world organizational challenges becomes critical. The results suggest that models that excel in producing polished responses may still fail in practical leadership tasks, such as information retrieval, trust maintenance, and decision execution. For organizations, this underscores the importance of evaluating AI not just on correctness but on its ability to manage consequences, prioritize effectively, and uphold trust under pressure. The experiment highlights that true leadership in AI involves managing complex, multi-faceted tasks that impact business outcomes, not just generating plausible text.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background of the Firmulate Management Benchmark

The Firmulate platform was launched to address a longstanding gap in AI evaluation: how models perform in managing real-world, high-stakes scenarios. Traditional benchmarks focus on technical output, coding, or conversational quality, but they do not assess how models handle crises, prioritize tasks, or maintain organizational trust. The live experiment in July 2026 simulated a small company’s worst week, with real money mechanics, versioned decisions, and a transparent process that allows for detailed analysis of model behavior over days. The scenario involved customer negotiations, crisis resolution, and internal management, providing a realistic testbed for AI management capabilities. The results reveal that, while models can recognize crises and resist manipulation, their ability to complete critical tasks and maintain trust varies significantly, exposing important gaps in current AI development.

“This experiment demonstrates that management quality, not just chat performance, should be a new benchmark for AI systems. The ability to manage consequences, prioritize correctly, and uphold trust is what truly distinguishes effective AI leadership.”

— Thorsten Meyer, Lead Developer at Firmulate

Amazon

organizational crisis management AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unresolved Questions About AI Management Performance

While the initial results are promising, it remains unclear how well these models will perform in larger, more complex organizations or in real-time operational environments. The experiment was conducted in a controlled, simulated setting with specific scenarios; whether these findings generalize to actual business contexts is still uncertain. Additionally, the impact of further training, customization, or integration with existing systems on management performance has not yet been established. The long-term reliability of models in maintaining trust and executing management tasks under sustained pressure also remains to be seen, as the current benchmark covers only a limited timeframe.

Amazon

AI decision-making evaluation platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for AI Management Benchmarking

Moving forward, firms are expected to expand these benchmarks to include more diverse scenarios and larger organizations, testing models’ scalability and robustness. Researchers plan to analyze how models adapt over extended periods and across different industries. Companies interested in deploying AI for management will likely use these benchmarks to evaluate and select models based on their ability to handle real-world consequences. Further development may include integrating human oversight and measuring how AI collaborates with human managers to improve decision quality and trust. The ongoing evolution of these benchmarks aims to establish management capability as a core criterion for AI readiness in organizational roles.

Amazon

AI leadership benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What does the new AI leaderboard measure?

The leaderboard assesses AI models’ ability to manage a simulated company’s crises, make decisions, uphold trust, and complete organizational tasks under pressure, highlighting management skills beyond traditional benchmarks.

Why is management performance important in AI evaluation?

Management performance is critical because AI systems are increasingly tasked with making complex decisions that impact business outcomes, trust, and reputation. Effective management involves not just producing correct answers but handling consequences and maintaining organizational integrity.

Can current models reliably manage real organizations?

While initial results are promising, it is still uncertain how well models will perform in larger, more complex real-world environments. Further testing and development are needed to confirm their reliability and robustness over time.

What are the main limitations of this benchmark?

The benchmark is based on simulated scenarios, which may not fully capture the complexity of real organizations. Its long-term applicability and performance in live settings remain to be seen.

What will happen next in AI management evaluation?

Expect expansion of benchmarks to include more diverse and larger organizational scenarios, with focus on scalability, robustness, and human-AI collaboration to ensure AI can effectively manage real-world consequences.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

California Hosts 34Th Annual Recovery Happens At State Capitol On September 2

California’s 34th Recovery Happens event took place at the State Capitol on September 2, highlighting ongoing recovery efforts and community support.

Target Field concessions workers set to begin strike Monday

Concessions workers at Target Field are set to begin a strike Monday, impacting game day operations and highlighting ongoing labor disputes.

First Horizon Bank Recognized With Trailblazer Award By United Way Of The Mid-South

First Horizon Bank received the Trailblazer Award from United Way of the Mid-South for community service and philanthropy, highlighting its local impact.

Board packet generator for HOA managers

A new board packet generator for HOA managers is being tested as a streamlined workflow for preparing board meetings, with initial validation underway.