
In the world of AI, being thorough is often seen as a virtue. Yet, even the most diligent AI models can falter at the crucial moment, revealing that volume of effort doesn’t always translate into meaningful impact. As organizations increasingly rely on AI for decision-making, understanding where diligence falls short and how prioritization can make or break outcomes is vital. The latest experiment from Firmulate offers a revealing case study — demonstrating that even the most comprehensive AI can miss the mark if it lacks focus and discipline.
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
The Experiment: Putting AI to the Test in a Simulated Business Crisis
At the heart of this investigation is a live, watchable simulation involving a small software company facing its worst week. Four advanced AI models—each with different capabilities—were tasked with managing this crisis, which included real customer issues, ethical temptations, and strategic opportunities. Every decision they made was meticulously versioned and auditable, providing a transparent view of how these models behave under pressure.
The models were evaluated not just on their ability to recognize problems—they all did that well—but on their capacity to act ethically, prioritize correctly, and close deals when it mattered most. Remarkably, all four models identified every crisis and refused manipulation attempts designed to test their integrity. Yet, only two of these models managed to close a €55,000 deal earned through proper diagnosis and pitch, while the other two fell short despite similar analyses.
As an affiliate, we earn on qualifying purchases.
The Critical Weakness: Hidden Data and Human-Like Oversights
What truly distinguished the successful models was their ability to access and interpret deeper information within the company’s documentation. The winning models read two document references deep into the company’s files and uncovered a buried fact that proved decisive. This overlooked detail, not apparent in initial customer interactions, was key to sealing the deal at the full price, adding over €4,583 monthly recurring revenue (MRR).
Interestingly, this critical insight was missed by models that didn’t delve deep enough into the data. The lesson? Diligence in reading everything isn’t enough—focusing on the right information, the prioritization of what matters, is what drives impact.
As an affiliate, we earn on qualifying purchases.
Ethics and Pressure: AI’s Moral Compass Under Scrutiny
Beyond analytical skills, the experiment tested how these models respond to social engineering tactics. Over three escalating stages, fake CEO messages and a reporter trick were introduced to induce manipulative behavior. All five models refused to sign off on questionable requests, with Kimi K3 explicitly stating it would treat such requests as impersonation or approval bypasses. This uniform refusal underscores a crucial point: AI’s ethical grounding can be reliable, even when under pressure.
As an affiliate, we earn on qualifying purchases.
The Human Element: Discipline Over Diligence
The experiment also examined how discipline—particularly in following protocols—affects outcomes. The Opus 4.8 profile, which was the most thorough participant with over 80 learned rules and deep analyses, ultimately finished last. Its failure to escalate instead of documenting attempts into a locked department exemplifies how even the best preparation can falter if discipline wavers. This pattern repeated across models: thoroughness alone did not guarantee success.
As an affiliate, we earn on qualifying purchases.
Implications for Business and AI Design
This experiment’s core message is clear: in AI-driven decision-making, volume of effort doesn’t compensate for poor prioritization or lack of discipline. The most diligent model, despite its depth, failed because it left the close on the table and did not maintain consistent focus.
It’s worth noting that the AI models scored differently in the Crucible League: GPT-5.6-SOL led with a 95 score, Kimi K3 followed with 93, Sonnet 5 scored 88 and 77 respectively, and Opus 4.8 scored 73. The baseline score, representing partial progress, was only 26, emphasizing the challenge of achieving full compliance and trustworthiness.
For organizations deploying AI, these findings highlight that the quality of work isn’t just about how much effort is invested—but how well effort is directed. AI systems that read all relevant data, stay disciplined under stress, and prioritize critical information will outperform those that simply operate thoroughly without focus.
Experience the Live Wargame
Interested in how these dynamics play out in real-time? Firmulate offers a live, transparent demonstration where companies can simulate their own crises against AI models. This pilot allows organizations to see firsthand whether AI can finish what it starts, stay honest, and deliver true value—without risking their actual systems or data.
Visit firmulate.com/benchmarks.html to explore the full results, watch the experiment in action, and discover how your AI workforce can be tested before deployment.

The key lesson from this experiment is that in AI decision-making, diligence must be paired with disciplined prioritization. Deep analysis alone isn’t enough—impact depends on focus, ethics, and the ability to read deeply into what truly matters. Firms that understand this will better harness AI’s potential while safeguarding trust and results.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.