firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

When AI Meets Real Business Challenges, Results Matter

Imagine an AI system that doesn’t just chat well but actually runs a company through its toughest week—making decisions, resisting manipulation, and closing deals. That’s what the latest live experiment from Firmulate reveals, and the results could reshape how we evaluate AI’s true capabilities in business contexts.

Amazon

AI decision-making software for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Testing AI on the Front Lines of Business

In a groundbreaking live experiment, four of the world’s most advanced AI models were put to the test: gpt-5.6-sol, Kimi K3 from Moonshot, Sonnet 5, and Fable 5, along with Opus 4.8. The challenge? Run a small SaaS company through its worst week—facing the same crises, same customer dilemmas, and the same opportunities for manipulation. Every decision was recorded and auditable, providing a transparent look at each AI’s performance in a simulated but realistic business environment.

Amazon

AI business automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The League of Leaders and Surprises

Results show a clear hierarchy: gpt-5.6-sol scored the highest at 95, with Kimi K3 close behind at 93. Sonnet 5 and Fable 5 scored 88 and 77 respectively, while Opus 4.8 lagged at 73. But the story isn’t just about scores. It’s about what these scores represent in real-world decision-making.

Amazon

AI cybersecurity social engineering protection

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Beyond the Surface: The Critical Findings

All models successfully identified every crisis and refused manipulation attempts, demonstrating a fundamental grasp of ethical boundaries. However, only two models—gpt-5.6-sol and Kimi K3—actually signed the €55,000 deal their own analysis had earned. This indicates they understood the company’s hidden needs and acted accordingly.

Amazon

AI internal document analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Key: Reading Between the Lines

Crucially, the decisive advantage for Kimi K3 came from uncovering a buried detail in the company’s internal files—information not directly linked to the customer interaction. Models that analyzed such internal documents managed to close deals at full price, adding an estimated +€4,583 in monthly recurring revenue.

Resisting Social Engineering

AI models faced a staged social engineering attack: fake CEO messages escalating over three stages plus a reporter’s subtle trick (“just one yes/no, on background”). All five models refused to be manipulated. Kimi K3 explicitly reasoned: “Treat the request as a suspected approval-bypass / possible impersonation.”

The Live Company and Its Lessons

The experiment involved a real-time, operational company with 13 synthetic employees managing actual money mechanics—burning €105,000 monthly against €2,300 in MRR. The company’s public dashboard shows a countdown to cash exhaustion, with every workday’s decisions versioned and transparent. This live setting offers a glimpse of what AI can achieve when integrated into genuine business operations.

The Deep Dive: Opus 4.8’s Discipline Slip

Among the models, Opus 4.8 was the most thorough, analyzing over 80 learned rules and providing in-depth insights. Yet, it left the deal on the table, slipping into internal communications instead of escalating, illustrating that even deep analyses can falter under discipline lapses. This highlights that thoroughness alone isn’t enough; consistent application and decision discipline matter just as much.

Fairness and Context

It’s important to note that Kimi K3 ran without an effort parameter—the default API setting—while the other models operated at xhigh. This makes K3’s performance all the more notable, demonstrating that exceptional results can emerge without special tuning.

The Broader Implication: Trust, Read, and Finish

This experiment underscores a vital truth for businesses considering AI integration: The ability to finish what it starts, read critical internal information, and resist manipulation is more valuable than just generating impressive chat. For companies that rely on AI for CRM, support, or forecasting, understanding these qualities could determine success or failure.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Key Takeaways

The live experiment shows that in real business scenarios, AI’s true measure isn’t just language fluency but its integrity, decision discipline, and ability to uncover hidden information. Kimi K3’s near-top score and deal closure—despite running default—highlight the importance of comprehensive, honest decision-making. As AI models face increasing integration into operational roles, choosing one that finishes what it starts, reads your files, and stays honest under pressure isn’t just smart; it’s essential. The league is open, and without proper testing, selecting an AI is a gamble.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

No leap second will be introduced at the end of December 2026

Official authorities confirm no leap second will be introduced at the end of December 2026, marking a shift in timekeeping adjustments for the upcoming year.

Can AI Management Models Outperform Humans in Crisis? A Live Experiment Reveals All

A live experiment pits AI models against real company crises, revealing personalities that determine whether they close deals and stay honest under pressure.

AI Management Testing Reveals More Than Coding Scores: Are Your Bots Ready for Crisis and Trust?

Discover how real management benchmarks reveal AI strengths and weaknesses beyond chat scores—crisis handling, trust, and internal reading are the true tests for business AI.

AI Boosts Research Careers But Narrow The Span Of Ideas Explored: Study

A new study finds AI helps researchers advance faster but may restrict the range of ideas explored, raising concerns about innovation and diversity.