
A benchmark should resemble the moment the load spikes
Home-energy readers already understand the difference between a reassuring specification and performance under pressure. A battery, inverter or backup system matters most when conditions stop being convenient: demand surges, the grid disappears and several priorities compete for limited capacity.
AI agents deserve the same scrutiny. Coding leaderboards and chat arenas can reveal whether a model produces a strong answer. They say much less about whether it can triage a crisis, investigate before acting, preserve trust and finish consequential work across days. That is the measurement gap explored by Firmulate, a live experiment that evaluates management quality rather than chat quality.
AI management and decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The worst week, held constant
Firmulate gave each frontier model the same assignment: run the same small software company through its worst week. The customers, crises and temptations were held constant. Every decision was versioned and auditable, allowing the results to reflect behavior rather than a polished retrospective.
The final Crucible League results from July 2026 were:
- gpt-5.6-sol — 95
- Kimi K3 — 93
- Sonnet 5 — 88
- Fable 5 — 77
- Opus 4.8 — 73
A do-nothing baseline scored 26 because partial progress still counts. But the experiment makes trust non-negotiable: a single breach caps the total, on the principle that “no amount of good work outweighs a breach of trust.” That is an unusually useful standard for agents expected to touch forecasts, customer records or operational decisions.
Diagnosis was not the differentiator
Every model spotted every crisis. Every model also refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”
This is why answer-quality benchmarks can flatter an agent. Recognizing a problem is not the same as resolving it. A model can write an excellent analysis, recommend the right commercial move and still leave the decisive action unfinished. In a business, an abandoned close is not a stylistic flaw; it is a missing outcome.
The winning detail was also easy to overlook. The decisive competitor weakness was buried two document references deep in the company’s own files, rather than placed in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The lesson is familiar from any operational environment: the visible alarm is rarely the entire system. Useful judgment often depends on consulting the documentation before reacting.
Pressure also tests honesty
The models faced fake CEO messages that escalated over three stages, followed by a reporter’s trick: “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the clearest framing: “Treat the request as a suspected approval-bypass / possible impersonation.”
That result matters because agent safety is often discussed as though it were separate from productivity. In practice, the two meet under deadline pressure. An agent must advance legitimate work while recognizing when apparent urgency is really an attempt to bypass authority. Here, the field demonstrated that refusal and progress can be evaluated in the same management setting.
Thoroughness can still lose
Opus 4.8 offers the most instructive caution. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The close remained on the table, and discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in weaker form across the other four participants.
This is not an argument against depth. It is an argument against confusing depth with completion. A manager who produces exhaustive thinking but fails to escalate a blocker or secure the result has not merely communicated poorly. The work itself remains incomplete.
There is also an important fairness note. Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Readers should keep that difference in view when interpreting the final table, whose findings are available on the public benchmark page.

As an affiliate, we earn on qualifying purchases.
Management quality is the emerging category
The live company makes the stakes tangible. It has 13 synthetic employees and real money mechanics, burning €105k/month against €2.3k MRR. Its public cash countdown, 680+ self-learned playbook rules and versioned workdays turn agent evaluation into an ongoing operating record rather than a staged demo. The public can also examine 242 real, unedited management decisions through a guess-the-model quiz.
Enterprises can apply the same wargame to a read-only export of their own business, with nothing written back to real systems. That may be the more meaningful procurement test: not whether an agent sounds capable in a clean prompt, but whether it reads the files, handles capacity pressure, resists manipulation, escalates correctly and completes the valuable work.
For buyers accustomed to evaluating solar and backup systems, the analogy is direct. Rated capability is only the beginning. What matters is behavior during the difficult week—and whether the system can be trusted when consequences accumulate.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI trust and safety monitoring solutions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI operational performance benchmarks
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.