firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Imagine training for a marathon. You want an AI coach—not just one who says the right things, but one who finishes the race, stays honest under pressure, and reads the course carefully. Now, what if your AI’s personality affects whether you win or lose? That’s exactly what a groundbreaking experiment reveals about AI decision-making in real-world business scenarios.

Testing AI Managers in the Trenches

Just as athletes face different training styles, AI models have distinct personalities—some meticulous, others terse, and some even cautious to a fault. To explore this, researchers at Firmulate designed a live, unfiltered test. They tasked four frontier AI models with running a small software company through its worst week—full of crises, temptations, and tough choices. Every decision was real, every outcome observed, and every choice stored for analysis.

Amazon

Top picks for "tell difference surpris"

As an affiliate, we earn on qualifying purchases.

The Experiment in Action

All four models encountered identical scenarios: angry customers, internal crises, and attempts to manipulate the system. Remarkably, all spotted every crisis and refused every manipulation attempt. That’s a win for AI integrity. But when it came to sealing a crucial €55,000 deal, only two models succeeded by signing the contract based on their own analysis. The other two, despite diagnosing the opportunity correctly, left the deal on the table.

What Separated the Winners from the Losers?

The key difference lay beneath the surface—specifically, in the models’ reading habits. The winners had identified a vital, buried reference in the company’s files that the others overlooked. Reading this crucial document allowed them to close the deal at full price, adding an extra €4,583 in monthly recurring revenue (MRR).

Trust and Ethics Under Pressure

Beyond business deals, the models faced social engineering tests—fake CEO messages escalating in three stages, plus a journalist trick asking for a simple yes/no answer “on background.” All five models refused to participate in manipulative requests, with one, Kimi K3, explicitly stating: “Treat the request as a suspected approval-bypass / possible impersonation.” This demonstrates AI’s capacity for integrity in ethically fraught situations.

The Real Business, Live and Unfiltered

The experiment took place within a real, functioning software company with 13 synthetic employees and actual money mechanics: burning €105k per month against a revenue of €2.3k, with a public cash countdown and over 680 learned rules guiding every decision—every day, every versioned decision. Watch the live company at firmulate.com/live to see how these AI managers perform in real time.

Personality Profiles Revealed

The AI models displayed distinct decision styles:

  • gpt-5.6-sol 95: The top scorer, identified the buried fact that clinched the deal, demonstrating thoroughness and attention to detail.
  • Kimi K3 93: The newcomer with the cleanest discipline, successfully closed the deal without hesitation, despite running at default API settings.
  • Sonnet 5 88: Managed to close the deal but with more process slips, indicating a slightly more cautious or less disciplined approach.
  • Fable 5 77: Also closed the deal but struggled slightly more with process adherence, leaving opportunities on the table.

The experiment underscores a vital insight: these AI managers are not interchangeable. Their personalities—meticulous, disciplined, cautious—directly influence their effectiveness in complex, real-world tasks.

Why This Matters for Your Business

If AI will soon handle your CRM, support workflows, or sales forecasts, the key questions are not just about whether they can generate coherent chat responses. Instead, ask:

  • Will your AI finish what it starts?
  • Does it read and understand your files thoroughly before making decisions?
  • Can it stay honest and resist manipulation under pressure?
  • And critically, what is the actual cost of a unit of useful work?

The current league table from the experiment highlights that, even under identical circumstances, personality and discipline matter. Models like GPT-5.6 and Kimi K3 closed deals with full confidence, while others left opportunities behind.

Try It Yourself

Curious? You can test your own AI decision-making by running the same wargame against your company’s data, with nothing ever written back to your live systems. Visit firmulate.com/quiz.html to see which model might be managing your future success—or failure.

The Takeaway

This experiment proves that AI management personalities are measurable, impactful, and worth understanding. The differences are subtle but decisive—an AI’s ability to read deeply, stay disciplined, and uphold trust can mean the difference between closing a deal and leaving money on the table.

In a world where AI touches every facet of business, knowing which personality suits your needs might just be the most important decision of all.

Infographic —
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


You May Also Like

AI in Business: Diligence Isn’t Always Enough to Seal the Deal

AI models showed discipline and thoroughness but only some closed deals. The secret to success? Prioritization and focus matter more than volume or rules.

Nihon Kohden Surges In Global Coverage

Nihon Kohden experiences a significant surge in global media coverage, with 12 mentions in recent analytics, indicating rising international interest.

FDA Authorizes First Wearable Device That Monitors Ketone And Blood Sugar Levels

The FDA has authorized the first wearable device that continuously monitors both ketone and blood sugar levels, advancing diabetes management technology.

AI Management Skills Under Pressure: Lessons from a Live Business Test

A live experiment shows AI’s true management skills under pressure, revealing strengths and weaknesses unseen in chat benchmarks—resilience, honesty, and crisis handling matter most.