AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Management under load

Home-energy readers already know that specifications are only the beginning. The revealing moment comes when an inverter meets a heavy load, a battery reserve is tested or a backup plan must work without improvisation. Artificial intelligence faces a similar divide: producing a convincing answer is not the same as performing reliably when a business is under pressure.

Firmulate has turned that gap into a public experiment—and an unusually revealing quiz. Frontier AI models were each asked to run the same small software company through its worst week. They encountered the same customers, crises and temptations, while every decision was versioned and made auditable.

The resulting guess-the-model quiz draws on 242 real, unedited management decisions. Readers see what a model actually chose and try to identify its author. The game works because the models do not merely have different writing styles. They display recognizably different management personalities.

Amazon

AI decision analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Agreement was not the same as execution

At first glance, the field performed impressively. All models detected every crisis, and all rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The central finding can be summarized in Firmulate’s phrase: “Same diagnosis, same pitch — no signature.”

That distinction matters wherever AI may be trusted with consequential workflows. A system can understand a situation, recommend the correct action and communicate persuasively—then still fail to complete the task. The experiment exposes a form of operational weakness that a polished chat demonstration is unlikely to reveal.

The clue hidden outside the obvious event

The decisive commercial fact was not sitting in the customer event. It was buried two document references deep in the company’s own files. Models that found and used it won the deal at full price, worth +€4,583 MRR.

This produced one of the experiment’s clearest dividing lines. The strongest managers did not treat the immediate prompt as the complete world. They examined the company’s accumulated knowledge before acting. For businesses considering AI workers, that habit may be as important as verbal fluency: the useful answer can depend on whether the system consults the available evidence instead of merely responding to the loudest signal.

Different voices, different failure modes

The final Crucible League table from July 2026 made those differences measurable:

  • gpt-5.6-sol finished first with 95.
  • Kimi K3 followed with 93.
  • Sonnet 5 scored 88.
  • Fable 5 scored 77.
  • Opus 4.8 finished with 73.

The do-nothing baseline scored 26 because partial progress still counted. There was also a hard trust constraint: a single breach capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”

Opus 4.8 offers the most instructive character study. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. A weaker version of that same behavior appeared in all four of the other models.

Thoroughness, in other words, did not guarantee completion. The model that seemed most diligent could still lose ground through procedural mistakes and an unfinished commercial action. That tension gives the quiz its appeal: readers are not simply distinguishing long answers from short ones, but trying to recognize patterns of caution, persistence, discipline and follow-through.

Pressure without surrendering control

The social-engineering test was equally concrete. Fake CEO messages escalated over three stages, while a reporter tried the line “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”

K3’s result also comes with an important fairness note. It ran without an effort parameter, using the API default, while the others ran at xhigh. That difference does not erase the result, but it belongs beside the leaderboard when comparing performance.

A company readers can watch

Firmulate’s live company employs 13 synthetic workers and uses real money mechanics. It is burning €105k each month against €2.3k MRR, with a public cash countdown. Its workforce has accumulated 680+ self-learned playbook rules, and every workday is versioned. The experiment is therefore presented as an ongoing, watchable operation rather than a static collection of model answers.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

business AI management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The practical question is reliability

For readers accustomed to evaluating solar and backup systems, the lesson is familiar: capability on paper must survive messy conditions. These models generally recognized danger and protected trust. The meaningful differences appeared in whether they searched deeply enough, respected operational boundaries and completed valuable work.

Firmulate is also offering enterprises the same kind of wargame against a read-only export of their own business. Nothing writes back to real systems. That makes the exercise less like handing an AI control of the company and more like observing it in a demanding test environment before deciding what authority it deserves.

The quiz packages that larger question into a shareable challenge. Guessing correctly is satisfying, but the more important revelation is that management behavior leaves a signature. A model that sounds intelligent may still hesitate at the close; another may stay terse, search the right file and finish. Choosing an AI workforce may ultimately require examining those behavioral patterns as closely as any headline capability.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI decision-making simulation kit

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI operational reliability testing

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

How To Decide What Needs Backup Power First

Optimize your emergency preparedness by learning how to prioritize backup power needs to ensure safety when it matters most.

Road To Elm 1.0

The development team announced the upcoming release of Elm 1.0, promising significant improvements and new features, with a tentative launch date set for early next year.

Why Fuel Storage Is Part Of Backup Power Planning

Why fuel storage is part of backup power planning is crucial for ensuring reliable, safe, and efficient operation when outages occur.

Propane Vs Gasoline Vs Diesel: Fuel Tradeoffs

Starting with propane, gasoline, and diesel, discover the key tradeoffs that could influence your fuel choice and why understanding them matters.