
Imagine testing a new AI system with the simplest task—nothing complex, just a baseline. Surprisingly, it scores a 26 out of 100, not zero. This reveals much about how we evaluate AI today, especially when trust and performance are on the line. In a world increasingly reliant on AI to handle real business crises, understanding what a minimal effort looks like—and how it’s measured—is essential for anyone who cares about digital reliability and integrity.
Prime made for students and young adults
- Fast, free delivery for dorm and study essentials
- Prime Video and Amazon Music included
- Member-only deals
Why a Do-Nothing Baseline Gets a Score of 26
At first glance, it might seem obvious that if an AI does nothing, it should score zero. But in the recent Firmulate experiment—a real, publicly observable test—the baseline AI scored 26 points. This isn’t a glitch or a mistake. It’s a reflection of how the scoring system is designed to account for partial progress and the inherent complexity of real-world decision-making.
In the experiment, four frontier AI models ran the same simulated week of crises in a small software company. These included customer issues, internal crises, and manipulative social engineering attempts. Each model was scored based on their ability to identify problems, resist manipulation, and ultimately close a deal. The key takeaway? Even the most passive or conservative models made some headway, earning partial credit. But more importantly, the scoring caps the total at 26 if any breach of trust occurs, ensuring that no amount of good work can outweigh dishonest behavior.
AI performance benchmarking tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Depth of the Methodology: How Progress is Measured
The scoring approach used by Firmulate emphasizes transparency and honesty. Each decision the AI models made was versioned and auditable. That way, researchers and business leaders could verify exactly how the models responded to each crisis, whether they read the files, how they handled manipulative requests, or whether they followed their own rules.
For example, models faced a social engineering test involving fake CEO messages escalating in complexity. All models refused these attempts—showing high trustworthiness. Yet, the real differentiator was in the details: the models that read certain internal documents were able to secure the deal at full price, while others left the opportunity on the table, costing the simulated company €4,583 in monthly recurring revenue.
AI trustworthiness testing software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why Full Trust Is Hard and How It’s Measured
One of the experiment’s key revelations is that a single breach of trust caps the total score, regardless of other achievements. This means that honesty is not just a bonus or a checkbox—it’s a foundation. A model that cheats, even once, can’t surpass a certain performance threshold. This approach mirrors real-world expectations: trustworthiness isn’t optional when dealing with sensitive or critical business decisions.
As an affiliate, we earn on qualifying purchases.
What the Results Say About AI Readiness for Business
The experiment shows that models like Kimi K3 and Sonnet 5 not only spotted crises but also avoided manipulative tactics, with K3 being the most disciplined. Interestingly, the most thorough participant, Opus 4.8, scored the lowest overall because it slipped in closing the deal and failed to escalate some issues properly. These nuanced results highlight that capability isn’t just about problem detection but also about disciplined process execution and honesty under pressure.
For business leaders, the takeaway is clear: performance isn’t solely about generating convincing output or chat quality. It’s about whether the AI can finish what it starts, read and understand internal files, resist temptation to cheat, and stay consistent in its decision-making. Only then can it serve as a reliable partner in managing real crises and making critical decisions.
As an affiliate, we earn on qualifying purchases.
What’s Next? Using Live Wargames to Test Your AI Workforce
Firmulate offers a practical way to assess your AI models before deployment through live wargames—simulating your own company’s worst week without risking real systems. These tests are transparent, versioned, and watchable online. They provide a clear picture of whether your AI can handle genuine business pressures or just excel in polished demos.
As the industry advances, benchmarks like the Crucible League serve as honest, hard tests—highlighting both strengths and weaknesses. And the key lesson? A do-nothing baseline isn’t zero; it’s a starting point, revealing that even the most conservative AI efforts have room for improvement, especially in trust and thoroughness.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
