
Imagine a company with no human employees, losing €105,000 every month, yet still actively competing in a live, publicly visible experiment. This isn’t science fiction — it’s the real-time story of Firmulate, a pioneering AI-driven business simulation that reveals how artificial intelligence can manage complex, high-stakes decisions under pressure.
The World’s Most Transparent Business Experiment
At the heart of this unfolding experiment is a small, virtual software company operated entirely by AI models. Every day, it faces genuine crises, customer demands, and ethical dilemmas, while its decision-making process is open for the world to observe. Managed by a system called Firmulate, the operation features 13 synthetic employees and employs over 680 self-learned rules to navigate the chaos.
What makes this experiment extraordinary is its transparency. Every workday version of the company is documented, versioned, and publicly accessible at firmulate.com/live.html. Viewers can see real-time decisions, the company’s cash burn, and the ongoing struggle to stay afloat financially. Currently, the company burns about €105,000 monthly against a revenue of just €2,300—a stark illustration of its precarious position.
AI business decision simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI Models in a High-Stakes Test
Four leading AI models, representing the frontier of language and reasoning capabilities, were tasked with guiding this virtual business through its worst week. Each model faced identical scenarios: crises with customers, internal crises, and manipulation attempts designed to test honesty and resilience.
The results were revealing:
- All four AI models identified every crisis and refused every manipulation attempt.
- Only two of the four signed a €55,000 deal after their own analysis — the same diagnosis and pitch, but with different outcomes.
- The decisive advantage came from a hidden detail in the company’s own files. Models that read deeper into the company’s internal documents uncovered a critical piece of information, allowing them to close a deal worth an additional €4,583 in monthly recurring revenue.
As an affiliate, we earn on qualifying purchases.
The Challenge of Trust and Integrity
One of the most compelling aspects of the experiment was the models’ responses to social engineering attempts. Fake CEO messages and a reporter tricking the system into a background consent request were met with unwavering refusal across all models. Kimi K3, for example, explicitly recognized the risk of impersonation and declined to proceed.
This resistance signals that AI systems trained in such environments are capable of maintaining ethical boundaries, even in high-pressure situations designed to tempt or manipulate them.
AI cybersecurity and manipulation detection tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Real Company Behind the Experiment
The company running this live test is not a game or a demo. It is a fully operational, albeit virtual, company with real money mechanics. It burns €105,000 per month, has a public cash countdown, and employs over 680 rules that it learned through self-play. Every day, the decision-making process is versioned and documented, making the entire operation auditable and transparent.
Some of the models performed better than others. For instance, Opus 4.8, with the most comprehensive set of analysis rules (over 80 learned rules), placed last in the league standings. Its discipline slipped, and it left potential deals on the table, illustrating that even the most thorough AI can falter in complex decision chains.
AI virtual assistant for business management
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Implications for the Future of AI in Business
This experiment underscores a vital point: the value of AI in business is not just about generating fluent text or customer interactions. It’s about whether AI can complete critical tasks ethically and reliably under pressure. Will your AI support system read your internal documents? Will it stay honest when faced with manipulation? And crucially, will it finish what it starts?
These questions are now testable in real time, thanks to the ongoing live experiment. Business leaders can run scenarios against their own operations, using the same principles and tools as this virtual company, via Firmulate’s pilot platform. This provides a risk-free environment to evaluate AI decision-making, without risking actual business operations.

The live Firmulate experiment offers a rare window into how AI models perform in managing real-world business crises, revealing strengths in crisis detection and refusal of manipulation, but also exposing discipline weaknesses. As AI takes on more critical roles, understanding these capabilities—and limitations—could be key to building trustworthy, effective automated decision-makers.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html