
Imagine a high-stakes workout—your trainer pushes you to your limits, testing not just your strength but your ability to stay honest, focused, and resilient under stress. Now, what if your trainer was an AI, running a small business through its toughest week yet? The results could change how we judge AI’s true readiness for real-world management.
Beyond Chatbot Glibness: Measuring True Management Competence
In the world of AI, much of the public dialogue revolves around how well models generate language—how convincingly they chat, answer questions, or simulate human-like interaction. However, in a recent live experiment conducted by Firmulate, the emphasis was on something far more critical: whether AI can effectively manage a small software company during its worst week. This wasn’t a test of linguistic finesse but a rigorous assessment of management quality under pressure.
The Setup: Real Crises, Real Money, Real Decisions
The experiment involved four leading AI models running an actual, functioning business—complete with real customer interactions, cash flow mechanics, and operational rules. Every decision was recorded and auditable, and the models faced identical scenarios, including crises, temptations, and manipulations. Their goal? To diagnose problems, make decisions, and close deals, just as a human manager would.
Key Findings: AI’s Capacity to Handle Crises and Maintain Honesty
Remarkably, all four models identified every crisis and refused every attempt at manipulation, indicating a strong grasp of compliance and integrity. Yet, only two managed to close the deal worth €55,000, which their own analysis had earned them. The twist? The decisive advantage was a buried piece of information—something that only the models trained to read deep into documents could uncover. Reading two document references deep in the company’s files was what made the difference, earning a full-month recurring revenue increase of over €4,583.
Behavior Under Stress: Discipline and Trust
One model, Opus 4.8, with the most thorough rule set, performed admirably in analysis but faltered at the critical moment, leaving the deal on the table and slipping into disciplinary silos instead of escalating issues appropriately. The models also faced social engineering tests—fake CEO messages and a reporter survey—where all refused to be manipulated, showing they could handle deception without compromising integrity.
Real Business, Real Money—Not Just a Game
The live company operated with 13 synthetic employees, burning €105k monthly against only €2.3k in monthly recurring revenue, with a public cash countdown. Every day, the models made decisions that impacted real money, and their performance was visible for all to see. This ongoing experiment, available at firmulate.com, is not a simulation—it’s a real-time test of management skills, not chat or language ability.
The True Benchmark: Management, Not Language
The experiment highlights a crucial gap: current benchmarks and leaderboards focus on answer quality and conversational ability, but they miss fundamental management qualities like reading comprehension, integrity, resilience under stress, and accountability. The leaderboard scores reflect superficial competence—gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, and Opus 4.8 at 73—yet, these numbers don’t tell the whole story of whether an AI can truly run a business or handle crises without slipping up.
The Broader Implication: AI as a Management Tool
In real-world applications—customer support, CRM, forecasting—the question isn’t just whether an AI writes well. It’s whether it can finish what it starts, read critical documents, stay honest under pressure, and adapt when crises strike. The live experiment underscores that true management ability involves resilience, integrity, and deep understanding—qualities that current chat benchmarks cannot capture.
Takeaway: Prepare Your AI Workforce for the Real Tests
As companies consider integrating AI into their management and operational workflows, the lesson is clear: testing with real crises and operational scenarios reveals strengths and weaknesses invisible in chat demos. Firmulate’s live wargame platform allows enterprises to simulate their own business challenges without risking real systems, ensuring AI models are ready for the demands of the real world.
In the end, the challenge isn’t just building models that talk well—it’s developing AI that manages responsibly, resists manipulation, and sustains performance under pressure. The future of AI in business will depend on this deeper measure of management quality, not just conversational skill.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html