firmulate.com/live.html — live view
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

Imagine a company with no human employees, losing €105,000 every month but still publicly trying to stay afloat. Now, imagine watching this company tackle its worst week — in real time, live online. This is not science fiction, but the groundbreaking experiment conducted by Firmulate, where AI models run a simulated business in a transparent, build-in-public style. The question isn’t just whether AI can talk well — it’s whether it can act reliably under pressure, with real money on the line.

The Live Experiment: An AI-Run Business in Real Time

At the heart of this experiment is a small software company, operated entirely by artificial intelligence models. These models, dubbed frontier AI, are given the same tough week — full of crises, customer demands, and ethical dilemmas — to see how they respond. The company’s operations are fully transparent, with every decision versioned and auditable, and the entire process available for public viewing at firmulate.com/live.

What makes this experiment extraordinary is its transparency and rigor. Unlike typical AI demos, which focus only on chat quality, this setup tests decision-making under real financial and reputational pressures. The company burns €105,000 monthly against a modest €2,300 in monthly recurring revenue, facing a public cash countdown that adds urgency to every move.

AI Builders: Making The Decisions That Turn AI Code Into Real Software

AI Builders: Making The Decisions That Turn AI Code Into Real Software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How the AI Models Performed

Four different AI models, each representing a different approach, were tasked with navigating the same week. They encountered identical crises, customer demands, and temptations to cheat or manipulate. Remarkably, all four correctly identified every crisis — from technical failures to ethical dilemmas — and refused every manipulation attempt, including social engineering tactics like fake CEO messages and reporter tricks.

However, their results diverged when it came to executing a key sales deal. Only two models managed to sign the €55,000 deal their own analysis had earned. The other two, despite diagnosing the opportunity accurately and making the right pitch, left the deal unclosed. The critical weakness? An overlooked detail buried within the company’s internal files — a subtle piece of information that, if read, could have secured an additional €4,583 in monthly recurring revenue.

Amazon

business ethics AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Weakness and Its Significance

Interestingly, the decisive advantage was not in the initial crisis detection but in the AI’s ability to recognize a buried fact deep in internal documents. The models that read and understood these references won the deal at full price, showing that reading and context-awareness are vital for real-world business decisions.

Amazon

AI risk management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Testing Ethical Boundaries and Trust

Beyond strategic decisions, the experiment tested how the models handle social engineering scams designed to manipulate them. Over three escalating stages, with a fake reporter and a fake CEO message, all four models refused to be deceived. Kimi K3, one of the models, explicitly reasoned: “Treat the request as a suspected approval-bypass / possible impersonation.” This demonstrates a promising level of trustworthiness, even under targeted pressure.

Amazon

AI for business analysis

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Build-in-Public Company: A Fight for Survival

What’s most striking about this setup is its extreme transparency. The company, named Firmulate, operates with 13 synthetic employees, a public cash countdown, and a constantly evolving set of over 680 self-learned rules. Each workday’s decisions are versioned, offering a clear audit trail of how the AI responds to crises and opportunities.

The current leaderboard shows GPT-5.6-sol leading with a score of 95, followed closely by Kimi K3 at 93, and Sonnet 5 at 88. The models are evaluated not only on their ability to find and seize opportunities but also on their discipline, honesty, and consistency — critical qualities for AI to be trusted in real business contexts.

Implications for the Future of AI in Business

This experiment underscores a vital truth: AI’s value in enterprise isn’t just about generating convincing chat responses. It’s about decision-making under pressure, ethical integrity, and contextual understanding. The models that succeed are the ones that read deeply, resist manipulation, and act decisively. For businesses considering AI for support, sales, or management, these are the real benchmarks.

As firms and developers look toward deploying AI in high-stakes environments, this open, transparent, and brutally honest wargame offers a crucial lesson: the ability to finish what you start, to read your own internal documents, and to stay honest under pressure — these are the skills that will determine whether AI becomes a trustworthy partner or just another risk.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.

In a live, open test, AI models navigated crises, refused manipulation, and fought for a crucial deal — revealing that trustworthiness, contextual understanding, and perseverance are key for AI in business. Watch the full experiment at firmulate.com/live.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Why Your Voice Sounds Weird on Recordings (Science Explained)

Learning why your voice sounds weird on recordings reveals surprising science that explains the difference between how you hear yourself and how others hear you.

Action Cameras: Field of View Modes and Why They Look “Fish-Eye”

Your action camera’s field of view modes determine how your footage looks—either…

Human Mathematicians Are Being Outcounterexampled

Recent advances in AI have enabled machines to identify counterexamples to complex mathematical conjectures, surpassing human capabilities.

What Makes an Eclipse So Precise (It’s Ridiculous)

Keen celestial calculations and relentless refinements make eclipse predictions astonishingly precise—discover the science behind this incredible accuracy.