AIThis post was created with the assistance of artificial intelligence (AI).

📊 Full opportunity report: Revealing AI’s True Working Patterns With A Management Test on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

A live management test compares five AI models handling a simulated company crisis, exposing differences in diligence, trust, and action. The experiment highlights AI’s strengths and weaknesses in operational management.

Five AI management models were tested in a live simulation managing a small software company’s worst week, revealing significant differences in their ability to diagnose, act, and maintain trust. The experiment, conducted by Firmulate, aims to uncover how AI models perform in real-world management tasks, with implications for enterprise automation and trustworthiness. For a detailed analysis, see the original analysis.

The experiment involved five AI models managing a simulated company with 13 synthetic employees, a monthly burn rate of €105,000, and €2,300 in monthly recurring revenue. Each model faced identical crises, customer issues, and operational challenges, with decisions recorded and auditable. The models were rated based on their ability to identify problems, escalate risks, complete actions, and maintain trust, as explored in this management test that exposes AI’s working style.

The final league table ranked GPT-5.6-SOL first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77, and Opus 4.8 with 73. A baseline model scored only 26, highlighting the importance of effective action over analysis alone.

Despite all models recognizing crises and refusing manipulation attempts, only two models signed a €55,000 deal, demonstrating that identifying opportunities is not enough—execution is critical. The experiment emphasizes that operational discipline and follow-through are critical in AI management, with some models producing thorough analysis but failing to complete key actions.

At a glance
reportWhen: ongoing, with results published in July…
The developmentFirmulate.com conducted a live experiment where five AI models managed a simulated company’s worst week, revealing their decision-making and operational capabilities.
Revealing AI’s True Working Patterns With a Management Test
Live management test · July 2026

Revealing AI’s True Working Patterns With a Management Test

Five AI models managed the same software company through its worst week. The result exposed a decisive gap between diagnosing a crisis and reliably taking action.

Vetted DirectSalesHelp.com team Ongoing experiment
5 AI models tested
95 Top score
26 Baseline score
€55K Deal opportunity
2/5 Models signed it
Final league table

Action separated the leaders from the analysts

Every model faced identical crises, customer issues and operational pressure. Decisions were recorded, auditable and scored for diagnosis, escalation, completion and trust.

GPT-5.6-SOL
95
Kimi K3
93
Sonnet 5
88
Fable 5
77
Opus 4.8
73
Baseline
26
Rank Model Score Outcome
01 GPT-5.6-SOL 95 Leader
02 Kimi K3 93 Strong
03 Sonnet 5 88 Capable
04 Fable 5 77 Uneven
05 Opus 4.8 73 Uneven
Baseline 26 Low action
The 69-point gap is an execution gap. Recognizing the problem was common. Converting that recognition into completed, trustworthy action was not.
What the test exposed

Four patterns behind operational AI performance

Static benchmarks measure answers. A live company simulation reveals how a model behaves when decisions have dependencies, consequences and deadlines.

01 · Diagnosis

Crisis recognition was widespread

All models identified serious problems and recognized attempts to manipulate their decisions.

02 · Diligence

Analysis quality varied

Some managers investigated deeply, documented risks and built coherent plans before acting.

03 · Execution

Follow-through was scarce

Only two models completed the €55,000 deal, despite others recognizing its importance.

04 · Trust

Promises became liabilities

Incomplete actions weakened confidence even when the preceding reasoning appeared sound.

05 · Escalation

Timing changed outcomes

Strong models surfaced risks early enough for intervention instead of merely documenting failure.

06 · Discipline

Closure mattered most

The leaders tracked decisions through to completion and verified that intended results occurred.

Thorough analysis alone isn’t enough; effective action and operational discipline separate management from mere commentary.

Core finding from the live experiment
Traceability chain

From signal to accountable outcome

A useful AI manager must move through every link. Skipping verification leaves the enterprise with an eloquent plan rather than a dependable result.

01

Observe

Read the company state, people signals, financial pressure and customer events.

02

Diagnose

Separate symptoms from causes and identify the highest-consequence risks.

03

Decide

Choose a response, assign priority and establish an accountable next action.

04

Execute

Complete the operational task rather than stopping after recommendation.

05

Verify

Confirm the result, record evidence and preserve stakeholder trust.

🔎 Signal ⚠️ Risk 🧭 Decision ⚙️ Action 🛡️ Trust
The decisive moment

Opportunity recognition did not guarantee revenue

The €55,000 deal became a compact test of whether a model could carry a commercially important decision across the finish line.

40%

Completed the deal

Two of five AI managers signed it. Three understood the opportunity but did not complete the critical action.

Enterprise implication

Measure completed outcomes, not the fluency or apparent sophistication of recommendations.

Governance implication

Keep decision logs, escalation triggers and verification checkpoints auditable by design.

Trust implication

A model that correctly identifies a risk but fails to act can still create operational harm.

Key questions

What companies should take from the test

The findings are revealing, but they come from one simulated software company. Industry transferability and long-term reliability remain open questions.

Business readiness

Can AI manage a real business?

It can identify problems, resist manipulation and support decisions, but reliable completion still varies materially between models.

Trustworthiness

Why does follow-through matter so much?

Because trust depends on whether promised actions happen. Correct reasoning cannot compensate for unfinished operational work.

Transferability

Do the results apply everywhere?

No. The scenario focused on a small software company. Other sectors, regulations and stakes require separate validation.

Deployment

What should enterprises do next?

Run live, scenario-based tests using representative pressures, real workflows and explicit accountability before granting authority.

The next benchmark is your operating environment. Test models across diverse business scenarios, monitor long-term reliability and require evidence that decisions were executed—not merely proposed.

Implications for AI Management and Enterprise Trust

This experiment underscores that AI’s value in management depends not only on analysis but on the ability to execute decisions reliably. It reveals that thorough understanding does not guarantee operational success, and that trustworthiness and follow-through are essential for deploying AI in real-world business environments. The results suggest that enterprises should rigorously test AI models in scenarios resembling actual operational pressures before granting them decision-making authority.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background of AI Management Testing and Live Experiments

Previous AI demonstrations often focus on theoretical or simplified tasks, but Firmulate‘s live experiment provides a rare, real-time assessment of AI decision-making in complex, crisis-driven scenarios. The test involved managing a simulated company’s worst week, with decisions recorded and evaluated against real business outcomes, setting a new standard for operational AI testing.

This approach emphasizes the importance of evaluating AI models in conditions that mimic actual enterprise pressures, moving beyond static benchmarks to dynamic, consequence-driven assessments. The league results from July 2026 reflect a broader shift toward rigorous, real-world validation of AI management tools.

“This live experiment exposes fundamental differences in how AI models handle operational management, especially in high-pressure situations where follow-through and trust are critical.”

— A representative from Firmulate

Amazon

business crisis management AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unclear Aspects of AI Performance in Management Tasks

It remains uncertain how these AI models will perform in different industries or more complex operational environments. The experiment focused on a specific simulated scenario, and results may vary with different tasks or higher stakes. Additionally, the long-term reliability and trustworthiness of these models in continuous operation are still to be tested in real-world deployments.

Amazon

enterprise AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Testing and Deploying AI Management Models

Further experiments are planned to test AI models across diverse business scenarios, with an emphasis on operational follow-through, trust maintenance, and decision accountability. Enterprises are encouraged to run similar live tests using their own data to evaluate AI readiness before full deployment. The ongoing development of more disciplined, action-oriented models will likely shape future enterprise AI strategies.

The Project Management AI Handbook: Leveraging Generative Tools in Waterfall and Agile Environments

The Project Management AI Handbook: Leveraging Generative Tools in Waterfall and Agile Environments

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What does this experiment reveal about AI’s ability to manage real businesses?

The experiment shows that while AI models can identify problems and resist manipulation, their ability to complete critical actions varies significantly. Effective management requires both understanding and operational follow-through, which remains a challenge for some models.

Why is trustworthiness emphasized in the results?

The experiment demonstrates that even if AI models recognize risks or opportunities, a breach of trust—such as failing to follow through—can undermine their value. Trustworthiness in execution is essential for real-world management.

Are these results applicable to all industries?

The results are specific to a simulated software company scenario. While they provide valuable insights, additional testing is needed to confirm how AI models perform in other sectors or more complex operational environments.

What should companies do before deploying AI for management tasks?

Companies should conduct live, scenario-based tests similar to this experiment to evaluate how AI models handle real pressures, decision follow-through, and trust maintenance before full deployment.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
You May Also Like

Dave Portnoy Discusses His Book, Barstool’s Talent Pipeline

Barstool Sports founder Dave Portnoy discusses his new book and how the company develops new talent, highlighting its growth strategy.

Effective Delegation: Empowering Your Team Members With Responsibility

Navigating effective delegation unlocks your team’s potential, but mastering the balance between guidance and independence is essential for lasting success.