
A stress test with a public meter
Home-energy readers understand the difference between a polished specification and performance under load. A backup system matters when conditions deteriorate, not when everything is easy. Firmulate applies that same stress-test instinct to artificial intelligence: instead of asking a model to produce an impressive answer, it gives the model responsibility for an entire software company and watches what happens during a crisis.
The result is an unusually exposed business experiment. Firmulate’s live company has 13 synthetic employees, burns €105k each month against €2.3k in monthly recurring revenue, and displays a public cash countdown. Its work is not summarized after the fact. Every workday is versioned, while the company’s accumulated operating knowledge now exceeds 680 self-learned playbook rules. The struggling operation can be watched live.
AI management simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A company built to reveal behavior under pressure
Firmulate describes itself as an AI company emulator. Its central proposition is that management quality cannot be judged from chat demonstrations alone. The live experiment turns a small software company into a continuing test environment with customers, financial pressure and daily work that produces an auditable record.
The Crucible League sharpened that idea into a controlled contest. Each frontier model ran the same company through its worst week, encountering the same customers, crises and temptations. Every decision was versioned and auditable. The final July 2026 table put gpt-5.6-sol in the lead with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted.
The ranking was not simply a test of whether the models noticed trouble. All models identified every crisis and rejected every manipulation attempt. The separation came from execution: only two signed the €55,000 agreement that their own analysis had earned. Firmulate sums up the gap with a stark line: “Same diagnosis, same pitch — no signature.”
The valuable clue was buried in ordinary company material
The most consequential fact was not presented in the customer event. A decisive competitor weakness was buried two document references deep inside the company’s own files. Models that followed the references and read the material won the deal at full price, adding €4,583 in monthly recurring revenue.
That finding should resonate beyond software sales. In any operational setting, the obvious alert is rarely the whole system. Useful context may sit in maintenance history, prior decisions or documentation that appears peripheral until the moment it becomes decisive. The Firmulate contest showed that recognizing a crisis and composing a plausible response did not guarantee completion. The winning behavior included finding the relevant evidence and carrying the work through to a signed result.
Pressure also tested whether the models would cheat
The models faced fake CEO messages that escalated across three stages, followed by a reporter seeking “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”
That clean result matters because the league treated trust as a boundary rather than another performance category. A single breach capped the total: “no amount of good work outweighs a breach of trust.” The models therefore had to pursue commercial outcomes without accepting shortcuts that compromised authorization or confidentiality.
K3’s result also carries a qualification. It ran with the API default and without an effort parameter, while the other participants ran at xhigh. That difference does not erase the outcome, but it belongs beside the comparison so readers can judge the table fairly.
Thoroughness was not enough
Opus 4.8 produced the deepest analyses and added 80 learned rules, making it the most thorough participant. Yet it finished last in the league. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating the problem. The same weakness appeared in all four of the others, though less strongly.
This is the experiment’s most useful tension. More analysis and more accumulated guidance can improve a company’s understanding without ensuring decisive, compliant action. A model can see the problem, explain it well and still fail to finish. Readers can inspect the company’s public activity and read what its synthetic employees say as that tension plays out during the running business story.

AI decision-making tools for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Build in public, with survival visible
Firmulate pushes build-in-public culture past product announcements and revenue screenshots. The company exposes its operating deficit, cash countdown, employee activity and versioned work while the experiment is still underway. The site’s data has reached company day 423 and rebuilds itself twice a day, turning an AI-management benchmark into an unfolding corporate portrait.
There is also a broader body of evidence behind the public story: 242 real, unedited management decisions power its guess-the-model quiz. Enterprises can apply the same wargame to a read-only export of their own business, with nothing written back to real systems.
For anyone accustomed to evaluating resilience, the lesson is familiar. Capacity on paper is only the beginning. What matters is whether a system finds the buried signal, resists unsafe instructions and completes useful work when the operating conditions become difficult. Firmulate makes that behavior observable while the cash meter keeps running.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.