firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

A management lesson hiding in a machine’s worst week

In education, the most revealing test is often not whether someone knows the answer, but whether they can use what they know under pressure. Firmulate’s experiment puts that question to AI: can a model run a company through a crisis, protect it from manipulation and still make the decision that keeps the business moving?

The answer is more complicated than a tidy benchmark score. The experiment offers a live, watchable look at AI management—and a case for testing it against the decisions a real organization may face. Readers can watch the experiment at Firmulate.

Same company, same difficult week

In the final Crucible League, run in July 2026, each frontier model faced the same small software company, customers, crises and temptations. Every decision was versioned and auditable. The results placed gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77, and Opus 4.8 fifth with 73. The do-nothing baseline scored 26. Partial progress counts, but one breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

On one measure, the models were unanimous: all spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The diagnosis was there; the close was not. As the experiment puts it: “Same diagnosis, same pitch — no signature.”

The clue was in the company’s own files

The decisive competitor weakness was buried two document references deep in the company’s files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. It is a practical lesson in research: a sound conclusion depends on finding and using relevant evidence, not just recognizing that a problem exists.

The pressure also came through social engineering. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

The leaderboard does not tell the whole story. Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet finished last. It left the deal unsigned and slipped on discipline, attempting writes in a locked department instead of escalating. The same weakness appeared, more weakly, in all four models. There is also a fairness caveat: K3 ran without an effort parameter, using the API default, while the others ran at xhigh.

A company you can watch

Firmulate’s live company has 13 synthetic employees and real money mechanics: monthly burn of €105k against €2.3k MRR, alongside a public cash countdown. It has learned more than 680 playbook rules, and every workday is versioned. A separate quiz uses 242 real, unedited management decisions to let readers guess which model made each call.

These details make the experiment more than a one-off comparison. Readers can follow decisions in a continuing simulation, then consider what such a test might reveal about their own organization. The company is synthetic; the business pressures are presented as a way to examine model behavior, not as a claim that an AI system has taken over a real business.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

From watching to a company-specific test

For an organization weighing AI agents in customer support, sales or operations, a general leaderboard can only go so far. Firmulate’s enterprise pilot applies the wargame to a read-only export of a company’s own business, then produces a board report with model rankings and weak points in its playbooks. Nothing writes back to real systems.

That turns the experiment into a practical question: how would different models handle your crises, evidence and approval boundaries? Explore the Firmulate pilot or contact contact@firmulate.com to discuss running the wargame against your business.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Heat Wave

A historic heat wave is affecting the UK, with temperatures surpassing previous records. Authorities warn of health risks and disruptions.

7.1 Earthquake In Japan

A 7.1 magnitude earthquake struck Japan, causing damage and evacuations. Authorities are assessing the impact, with no confirmed casualties reported yet.

Best Educational Science Kits For Students Compared

Compare popular educational science kits for students, focusing on content, price, age range, and value to help you pick the right kit for your learner.

Houston to Experience Rising Temperatures This Week Due to High-Pressure System

Temperatures in Houston are expected to increase this week as a high-pressure system settles over the region, bringing prolonged heat.