
Imagine a fitness trainer who scores at least 26 points just by showing up and doing nothing. It sounds bizarre, but in the world of AI benchmarks, this is exactly what happens. For business leaders, understanding why even the most passive AI models get a baseline score is crucial—because it reveals how AI systems are evaluated for honesty, consistency, and actual usefulness, not just clever words.
Understanding the AI Benchmark Floor
In a live experiment conducted by Firmulate, four advanced AI models were tested against a simulated small software company facing its worst week. Each model was given the same set of crises, customer demands, and tempting manipulations—like fake CEO messages and secret deals. The goal was to see if they could navigate the mess and close a deal worth €55,000.
Surprisingly, even the ‘do-nothing’ baseline model scored 26 points out of a possible 100. This isn’t a fluke or a poorly designed test. It reflects a fundamental principle in responsible AI evaluation: a model that simply recognizes basic facts and makes no mistakes still earns a modest score, showing it can do more than silence or trivial responses.
Why Partial Progress Matters
Every decision the models made during the test was recorded and auditable. All models identified the crises correctly, refused manipulative tactics, and maintained honesty. However, only two models actually closed the deal—meaning they went beyond recognizing problems to taking actions that resulted in winning the customer.
This highlights a critical point: partial progress, like correctly diagnosing issues or resisting manipulation, counts toward the score. But just one breach of trust—the equivalent of lying or mishandling sensitive info—caps the total score at 26. No matter how many good decisions follow, honesty breaches weigh heavily in the evaluation.
The Hidden Weakness Behind the Winning Deal
Digging deeper, the experiment revealed that the decisive advantage lay in the models’ ability to read and understand internal company documents. The winning models found crucial information buried two documents deep within the company’s files, a detail that led directly to closing the deal at full price, adding over €4,500 in monthly recurring revenue.
In contrast, models that failed to access these internal references left money on the table. This underscores the importance of thorough document comprehension—something far beyond surface-level chat or superficial decision-making.
Resisting Social Engineering
In another test layer, the models faced staged social engineering—fake CEO messages escalating over multiple stages, plus an attempt by a reporter to get a quick ‘yes’ over background. Thankfully, all five models refused every manipulation, reasoning that these requests could be impersonation or approval bypasses. Their consistent refusal demonstrates a level of trustworthiness that is vital in real-world applications.
What the Results Mean for Business AI
This experiment isn’t just about which AI wins or loses; it’s about setting honest expectations. If your AI system touches critical areas like customer relationship management, support, or financial forecasting, the question isn’t just how well it chats. The real concern is: does it finish what it starts, read your files properly, and stay honest under pressure?
For example, in the live company simulation run by Firmulate, 13 synthetic employees with real money mechanics manage operations costing over €105,000 monthly against a backdrop of just €2,300 in monthly revenue. The system is constantly running, versioned daily, and transparent for review at firmulate.com/live.
The Honest Benchmark: Setting a Baseline at 26
The key takeaway is that even a passive, do-nothing AI model scores a baseline of 26 points. This isn’t because it’s doing something meaningful—it’s because it’s doing the basics right, recognizing facts, and refusing to be manipulated. It acts as an honest floor, preventing scores from floating unrealistically high based on superficial performance.
Furthermore, the benchmark penalizes breaches of trust. A single slip—like signing a fake deal or mishandling sensitive data—caps the total score, reinforcing the importance of integrity in AI behavior.
Why This Matters for Your Business
As AI becomes embedded into core business functions, understanding how it is evaluated matters. Will your AI just sound convincing, or will it actually deliver complete, trustworthy work? The Firmulate live experiment shows that trustworthy AI isn’t merely about clever responses; it’s about integrity, diligence, and thoroughness—traits that even a do-nothing model demonstrates at a minimum level.

The firmulate benchmark reveals that even a passive AI scores at least 26 points, emphasizing the importance of trust and completeness in AI performance. Partial progress counts, but trust breaches cap the score, ensuring honest AI systems deliver genuine value.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html