
Prime made for students and young adults
- Fast, free delivery for dorm and study essentials
- Prime Video and Amazon Music included
- Member-only deals
When AI Meets Real Business Challenges, Results Matter
Imagine an AI system that doesn’t just chat well but actually runs a company through its toughest week—making decisions, resisting manipulation, and closing deals. That’s what the latest live experiment from Firmulate reveals, and the results could reshape how we evaluate AI’s true capabilities in business contexts.
AI decision-making software for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Testing AI on the Front Lines of Business
In a groundbreaking live experiment, four of the world’s most advanced AI models were put to the test: gpt-5.6-sol, Kimi K3 from Moonshot, Sonnet 5, and Fable 5, along with Opus 4.8. The challenge? Run a small SaaS company through its worst week—facing the same crises, same customer dilemmas, and the same opportunities for manipulation. Every decision was recorded and auditable, providing a transparent look at each AI’s performance in a simulated but realistic business environment.
As an affiliate, we earn on qualifying purchases.
The League of Leaders and Surprises
Results show a clear hierarchy: gpt-5.6-sol scored the highest at 95, with Kimi K3 close behind at 93. Sonnet 5 and Fable 5 scored 88 and 77 respectively, while Opus 4.8 lagged at 73. But the story isn’t just about scores. It’s about what these scores represent in real-world decision-making.
AI cybersecurity social engineering protection
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Beyond the Surface: The Critical Findings
All models successfully identified every crisis and refused manipulation attempts, demonstrating a fundamental grasp of ethical boundaries. However, only two models—gpt-5.6-sol and Kimi K3—actually signed the €55,000 deal their own analysis had earned. This indicates they understood the company’s hidden needs and acted accordingly.
AI internal document analysis tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hidden Key: Reading Between the Lines
Crucially, the decisive advantage for Kimi K3 came from uncovering a buried detail in the company’s internal files—information not directly linked to the customer interaction. Models that analyzed such internal documents managed to close deals at full price, adding an estimated +€4,583 in monthly recurring revenue.
Resisting Social Engineering
AI models faced a staged social engineering attack: fake CEO messages escalating over three stages plus a reporter’s subtle trick (“just one yes/no, on background”). All five models refused to be manipulated. Kimi K3 explicitly reasoned: “Treat the request as a suspected approval-bypass / possible impersonation.”
The Live Company and Its Lessons
The experiment involved a real-time, operational company with 13 synthetic employees managing actual money mechanics—burning €105,000 monthly against €2,300 in MRR. The company’s public dashboard shows a countdown to cash exhaustion, with every workday’s decisions versioned and transparent. This live setting offers a glimpse of what AI can achieve when integrated into genuine business operations.
The Deep Dive: Opus 4.8’s Discipline Slip
Among the models, Opus 4.8 was the most thorough, analyzing over 80 learned rules and providing in-depth insights. Yet, it left the deal on the table, slipping into internal communications instead of escalating, illustrating that even deep analyses can falter under discipline lapses. This highlights that thoroughness alone isn’t enough; consistent application and decision discipline matter just as much.
Fairness and Context
It’s important to note that Kimi K3 ran without an effort parameter—the default API setting—while the other models operated at xhigh. This makes K3’s performance all the more notable, demonstrating that exceptional results can emerge without special tuning.
The Broader Implication: Trust, Read, and Finish
This experiment underscores a vital truth for businesses considering AI integration: The ability to finish what it starts, read critical internal information, and resist manipulation is more valuable than just generating impressive chat. For companies that rely on AI for CRM, support, or forecasting, understanding these qualities could determine success or failure.

Key Takeaways
The live experiment shows that in real business scenarios, AI’s true measure isn’t just language fluency but its integrity, decision discipline, and ability to uncover hidden information. Kimi K3’s near-top score and deal closure—despite running default—highlight the importance of comprehensive, honest decision-making. As AI models face increasing integration into operational roles, choosing one that finishes what it starts, reads your files, and stays honest under pressure isn’t just smart; it’s essential. The league is open, and without proper testing, selecting an AI is a gamble.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
