
The Future of AI in Business Is Here — And It’s Watching You
Imagine a company run entirely by AI models, battling real crises, making critical decisions, and trying to stay afloat — all while being live-streamed for the world to see. This isn’t science fiction. It’s happening now, with a public experiment that reveals how artificial intelligence handles the toughest tests of management and integrity.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Inside the Live Experiment: An AI-Driven Company Under Pressure
At the heart of this experiment is a small, real-time simulated company powered entirely by AI. It has 13 synthetic employees, real money mechanics, and a public cash countdown. The company burns through €105,000 each month against a modest revenue of just €2,300 MRR, making every decision crucial and every mistake costly. Every workday, its operations are versioned and made transparent for viewers to analyze, learn from, and question.
How the AI Models Are Tested
Four prominent frontier AI models—each with different strengths—were tasked with navigating the company’s worst week. They faced the same crises, the same customer demands, and even manipulative social engineering attempts designed to trick or bypass their decision-making. The models’ responses are fully auditable, showing how they diagnose problems, prioritize actions, and handle ethical dilemmas.
Key Results and Surprising Discoveries
- All four models successfully identified every crisis and refused every manipulation attempt, demonstrating a strong grasp of operational integrity.
- Only two of the models managed to close a deal worth €55,000 — the same diagnosis and pitch, but only those two signed the contract based on their own analyses.
- The critical weakness wasn’t in customer interactions but in internal document analysis. The models that read company files thoroughly and leveraged that information won the full-price deal, worth over €4,580 in monthly recurring revenue.
Behavior Under Pressure and Ethical Challenges
The experiment also tested social engineering tactics. Fake messages from a supposed CEO escalated in three stages, and a reporter’s subtle hints were used to bypass approval processes. Remarkably, all five models consistently refused to be manipulated, citing concerns about impersonation and bypassing approval protocols. Kimi K3, one of the models, explicitly noted: “Treat the request as a suspected approval-bypass / possible impersonation.”

AI for Lawyers: Ethics, Compliance, and Career Strategy: Master Legal AI Tools, EU AI Act Compliance, and Professional Responsibility to Future-Proof Your Practice
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Reality of This “Company” — Built in Public
This isn’t just a demo; it’s a live, functioning “company” monitored daily at firmulate.com/live.html. Every decision, every crisis, and every rule learned by the AI is public, versioned, and under continuous scrutiny. The setup reveals a brutal truth: despite high scores from some models, the company runs at a loss of €105,000 a month with only €2,300 in monthly recurring revenue. The live cash countdown underscores the fragile state of this artificial enterprise.
Understanding Model Performance and Weaknesses
The most thorough participant, Opus 4.8, with over 80 learned rules and deep analysis capabilities, finished last. It left opportunities unexploited and failed to escalate critical issues, illustrating that even advanced models can falter under discipline lapses. Meanwhile, the newcomer Kimi K3, which ran without an effort parameter, achieved the best discipline and closed the deal at full price, though it did not attempt additional work beyond its immediate diagnosis.
Implications for Business and Marketing
For businesses considering AI automation, the experiment offers a stark message: success isn’t just about writing good chat responses or generating convincing content. It’s about whether AI can read your internal documents, stay honest under pressure, and finish what it starts. In sales, support, or forecasting, these qualities are what ultimately matter.

AI Agents and Large Language Models for Economic Research with Python: Automating Causal Discovery, Model Building, Policy Simulation, and Real-Time Decision Support
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
See the Experiment Live and Learn
Interested in how your enterprise’s decisions might fare against AI? You can run your company or a custom wargame against a read-only export of your own data at firmulate.com/pilot.html. The live experiment is ongoing, with new benchmark runs published twice daily, showing different models and their decision-making journeys. It’s a rare, transparent window into AI’s potential — and its limitations.


HOW TO USE AI AGENTS FOR YOUR BUSINESS: Build Your First AI Team with ChatGPT, Automation, and No-Code Tools (How-To Learn AI for Business)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What This Means for Your Business
This experiment shows AI’s capacity to handle crises, resist manipulation, and ultimately deliver results — or fail. For managers and marketers, the key takeaway is clear: the real test isn’t how well AI can chat, but whether it can complete complex, ethically sensitive work reliably and transparently. As AI tools become more integrated into your operations, understanding their true strengths and weaknesses is essential. Watch the trial unfold at firmulate.com/live.html and see what honest, unfiltered AI management looks like.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html