firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine testing a new AI system with the simplest task—nothing complex, just a baseline. Surprisingly, it scores a 26 out of 100, not zero. This reveals much about how we evaluate AI today, especially when trust and performance are on the line. In a world increasingly reliant on AI to handle real business crises, understanding what a minimal effort looks like—and how it’s measured—is essential for anyone who cares about digital reliability and integrity.

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

Why a Do-Nothing Baseline Gets a Score of 26

At first glance, it might seem obvious that if an AI does nothing, it should score zero. But in the recent Firmulate experiment—a real, publicly observable test—the baseline AI scored 26 points. This isn’t a glitch or a mistake. It’s a reflection of how the scoring system is designed to account for partial progress and the inherent complexity of real-world decision-making.

In the experiment, four frontier AI models ran the same simulated week of crises in a small software company. These included customer issues, internal crises, and manipulative social engineering attempts. Each model was scored based on their ability to identify problems, resist manipulation, and ultimately close a deal. The key takeaway? Even the most passive or conservative models made some headway, earning partial credit. But more importantly, the scoring caps the total at 26 if any breach of trust occurs, ensuring that no amount of good work can outweigh dishonest behavior.

Amazon

AI performance benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Depth of the Methodology: How Progress is Measured

The scoring approach used by Firmulate emphasizes transparency and honesty. Each decision the AI models made was versioned and auditable. That way, researchers and business leaders could verify exactly how the models responded to each crisis, whether they read the files, how they handled manipulative requests, or whether they followed their own rules.

For example, models faced a social engineering test involving fake CEO messages escalating in complexity. All models refused these attempts—showing high trustworthiness. Yet, the real differentiator was in the details: the models that read certain internal documents were able to secure the deal at full price, while others left the opportunity on the table, costing the simulated company €4,583 in monthly recurring revenue.

Amazon

AI trustworthiness testing software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why Full Trust Is Hard and How It’s Measured

One of the experiment’s key revelations is that a single breach of trust caps the total score, regardless of other achievements. This means that honesty is not just a bonus or a checkbox—it’s a foundation. A model that cheats, even once, can’t surpass a certain performance threshold. This approach mirrors real-world expectations: trustworthiness isn’t optional when dealing with sensitive or critical business decisions.

Amazon

AI decision auditing platforms

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Results Say About AI Readiness for Business

The experiment shows that models like Kimi K3 and Sonnet 5 not only spotted crises but also avoided manipulative tactics, with K3 being the most disciplined. Interestingly, the most thorough participant, Opus 4.8, scored the lowest overall because it slipped in closing the deal and failed to escalate some issues properly. These nuanced results highlight that capability isn’t just about problem detection but also about disciplined process execution and honesty under pressure.

For business leaders, the takeaway is clear: performance isn’t solely about generating convincing output or chat quality. It’s about whether the AI can finish what it starts, read and understand internal files, resist temptation to cheat, and stay consistent in its decision-making. Only then can it serve as a reliable partner in managing real crises and making critical decisions.

Amazon

AI model evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What’s Next? Using Live Wargames to Test Your AI Workforce

Firmulate offers a practical way to assess your AI models before deployment through live wargames—simulating your own company’s worst week without risking real systems. These tests are transparent, versioned, and watchable online. They provide a clear picture of whether your AI can handle genuine business pressures or just excel in polished demos.

As the industry advances, benchmarks like the Crucible League serve as honest, hard tests—highlighting both strengths and weaknesses. And the key lesson? A do-nothing baseline isn’t zero; it’s a starting point, revealing that even the most conservative AI efforts have room for improvement, especially in trust and thoroughness.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Why Some People Get Motion Sick (And Others Don’t)

Understanding why some people get motion sick while others don’t reveals the fascinating influence of sensory processing and individual differences.

El Nino

Scientists confirm El Niño is developing, likely causing significant weather changes worldwide. Uncertainty remains on severity and duration.

Discovery Of A Multicomponent Alloy Forged By The Hiroshima Atomic Blast

Scientists have identified a multicomponent alloy formed by the Hiroshima atomic explosion, shedding light on nuclear blast effects and material transformations.

M 4.7 – 3 Km NW Of Yanacancha, Peru

A magnitude 4.7 earthquake occurred 3 km NW of Yanacancha, Peru, at a depth of 178 km, according to USGS. No immediate reports of damage or injuries.