firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

When AI Meets Real Business Challenges, Results Matter

Imagine an AI system that doesn’t just chat well but actually runs a company through its toughest week—making decisions, resisting manipulation, and closing deals. That’s what the latest live experiment from Firmulate reveals, and the results could reshape how we evaluate AI’s true capabilities in business contexts.

Amazon

AI decision-making software for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Testing AI on the Front Lines of Business

In a groundbreaking live experiment, four of the world’s most advanced AI models were put to the test: gpt-5.6-sol, Kimi K3 from Moonshot, Sonnet 5, and Fable 5, along with Opus 4.8. The challenge? Run a small SaaS company through its worst week—facing the same crises, same customer dilemmas, and the same opportunities for manipulation. Every decision was recorded and auditable, providing a transparent look at each AI’s performance in a simulated but realistic business environment.

Amazon

AI business automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The League of Leaders and Surprises

Results show a clear hierarchy: gpt-5.6-sol scored the highest at 95, with Kimi K3 close behind at 93. Sonnet 5 and Fable 5 scored 88 and 77 respectively, while Opus 4.8 lagged at 73. But the story isn’t just about scores. It’s about what these scores represent in real-world decision-making.

Amazon

AI cybersecurity social engineering protection

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Beyond the Surface: The Critical Findings

All models successfully identified every crisis and refused manipulation attempts, demonstrating a fundamental grasp of ethical boundaries. However, only two models—gpt-5.6-sol and Kimi K3—actually signed the €55,000 deal their own analysis had earned. This indicates they understood the company’s hidden needs and acted accordingly.

Amazon

AI internal document analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Key: Reading Between the Lines

Crucially, the decisive advantage for Kimi K3 came from uncovering a buried detail in the company’s internal files—information not directly linked to the customer interaction. Models that analyzed such internal documents managed to close deals at full price, adding an estimated +€4,583 in monthly recurring revenue.

Resisting Social Engineering

AI models faced a staged social engineering attack: fake CEO messages escalating over three stages plus a reporter’s subtle trick (“just one yes/no, on background”). All five models refused to be manipulated. Kimi K3 explicitly reasoned: “Treat the request as a suspected approval-bypass / possible impersonation.”

The Live Company and Its Lessons

The experiment involved a real-time, operational company with 13 synthetic employees managing actual money mechanics—burning €105,000 monthly against €2,300 in MRR. The company’s public dashboard shows a countdown to cash exhaustion, with every workday’s decisions versioned and transparent. This live setting offers a glimpse of what AI can achieve when integrated into genuine business operations.

The Deep Dive: Opus 4.8’s Discipline Slip

Among the models, Opus 4.8 was the most thorough, analyzing over 80 learned rules and providing in-depth insights. Yet, it left the deal on the table, slipping into internal communications instead of escalating, illustrating that even deep analyses can falter under discipline lapses. This highlights that thoroughness alone isn’t enough; consistent application and decision discipline matter just as much.

Fairness and Context

It’s important to note that Kimi K3 ran without an effort parameter—the default API setting—while the other models operated at xhigh. This makes K3’s performance all the more notable, demonstrating that exceptional results can emerge without special tuning.

The Broader Implication: Trust, Read, and Finish

This experiment underscores a vital truth for businesses considering AI integration: The ability to finish what it starts, read critical internal information, and resist manipulation is more valuable than just generating impressive chat. For companies that rely on AI for CRM, support, or forecasting, understanding these qualities could determine success or failure.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Key Takeaways

The live experiment shows that in real business scenarios, AI’s true measure isn’t just language fluency but its integrity, decision discipline, and ability to uncover hidden information. Kimi K3’s near-top score and deal closure—despite running default—highlight the importance of comprehensive, honest decision-making. As AI models face increasing integration into operational roles, choosing one that finishes what it starts, reads your files, and stays honest under pressure isn’t just smart; it’s essential. The league is open, and without proper testing, selecting an AI is a gamble.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Will Minneapolis, MN Face Tornado Risk On September 15?

Uncertainty surrounds tornado risk in Minneapolis on September 15, with weather forecasts and models showing no definitive threat yet.

How To Sequence Your Own DNA At Home

A detailed guide on how individuals can now sequence their DNA at home using accessible tools, with confirmed methods and current limitations.

Texture Control: The Simple Variables Behind Crunch vs Chew

Prepare to master texture control by exploring simple variables that determine whether your food is crispy or tender—your perfect bite awaits.

Will The Temp In Austin Be Above 76.99° On Jul 12, 2026 At 5Am EDT?

A market-based prediction indicates a high interest in whether Austin’s temperature will be above 76.99°F at 5am EDT on July 12, 2026, but the event remains uncertain.