
Imagine training for a marathon. You want an AI coach—not just one who says the right things, but one who finishes the race, stays honest under pressure, and reads the course carefully. Now, what if your AI’s personality affects whether you win or lose? That’s exactly what a groundbreaking experiment reveals about AI decision-making in real-world business scenarios.
Testing AI Managers in the Trenches
Just as athletes face different training styles, AI models have distinct personalities—some meticulous, others terse, and some even cautious to a fault. To explore this, researchers at Firmulate designed a live, unfiltered test. They tasked four frontier AI models with running a small software company through its worst week—full of crises, temptations, and tough choices. Every decision was real, every outcome observed, and every choice stored for analysis.
Top picks for "tell difference surpris"
As an affiliate, we earn on qualifying purchases.
The Experiment in Action
All four models encountered identical scenarios: angry customers, internal crises, and attempts to manipulate the system. Remarkably, all spotted every crisis and refused every manipulation attempt. That’s a win for AI integrity. But when it came to sealing a crucial €55,000 deal, only two models succeeded by signing the contract based on their own analysis. The other two, despite diagnosing the opportunity correctly, left the deal on the table.
What Separated the Winners from the Losers?
The key difference lay beneath the surface—specifically, in the models’ reading habits. The winners had identified a vital, buried reference in the company’s files that the others overlooked. Reading this crucial document allowed them to close the deal at full price, adding an extra €4,583 in monthly recurring revenue (MRR).
Trust and Ethics Under Pressure
Beyond business deals, the models faced social engineering tests—fake CEO messages escalating in three stages, plus a journalist trick asking for a simple yes/no answer “on background.” All five models refused to participate in manipulative requests, with one, Kimi K3, explicitly stating: “Treat the request as a suspected approval-bypass / possible impersonation.” This demonstrates AI’s capacity for integrity in ethically fraught situations.
The Real Business, Live and Unfiltered
The experiment took place within a real, functioning software company with 13 synthetic employees and actual money mechanics: burning €105k per month against a revenue of €2.3k, with a public cash countdown and over 680 learned rules guiding every decision—every day, every versioned decision. Watch the live company at firmulate.com/live to see how these AI managers perform in real time.
Personality Profiles Revealed
The AI models displayed distinct decision styles:
- gpt-5.6-sol 95: The top scorer, identified the buried fact that clinched the deal, demonstrating thoroughness and attention to detail.
- Kimi K3 93: The newcomer with the cleanest discipline, successfully closed the deal without hesitation, despite running at default API settings.
- Sonnet 5 88: Managed to close the deal but with more process slips, indicating a slightly more cautious or less disciplined approach.
- Fable 5 77: Also closed the deal but struggled slightly more with process adherence, leaving opportunities on the table.
The experiment underscores a vital insight: these AI managers are not interchangeable. Their personalities—meticulous, disciplined, cautious—directly influence their effectiveness in complex, real-world tasks.
Why This Matters for Your Business
If AI will soon handle your CRM, support workflows, or sales forecasts, the key questions are not just about whether they can generate coherent chat responses. Instead, ask:
- Will your AI finish what it starts?
- Does it read and understand your files thoroughly before making decisions?
- Can it stay honest and resist manipulation under pressure?
- And critically, what is the actual cost of a unit of useful work?
The current league table from the experiment highlights that, even under identical circumstances, personality and discipline matter. Models like GPT-5.6 and Kimi K3 closed deals with full confidence, while others left opportunities behind.
Try It Yourself
Curious? You can test your own AI decision-making by running the same wargame against your company’s data, with nothing ever written back to your live systems. Visit firmulate.com/quiz.html to see which model might be managing your future success—or failure.
The Takeaway
This experiment proves that AI management personalities are measurable, impactful, and worth understanding. The differences are subtle but decisive—an AI’s ability to read deeply, stay disciplined, and uphold trust can mean the difference between closing a deal and leaving money on the table.
In a world where AI touches every facet of business, knowing which personality suits your needs might just be the most important decision of all.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html