AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

A benchmark should resemble the moment the load spikes

Home-energy readers already understand the difference between a reassuring specification and performance under pressure. A battery, inverter or backup system matters most when conditions stop being convenient: demand surges, the grid disappears and several priorities compete for limited capacity.

AI agents deserve the same scrutiny. Coding leaderboards and chat arenas can reveal whether a model produces a strong answer. They say much less about whether it can triage a crisis, investigate before acting, preserve trust and finish consequential work across days. That is the measurement gap explored by Firmulate, a live experiment that evaluates management quality rather than chat quality.

Amazon

AI management and decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The worst week, held constant

Firmulate gave each frontier model the same assignment: run the same small software company through its worst week. The customers, crises and temptations were held constant. Every decision was versioned and auditable, allowing the results to reflect behavior rather than a polished retrospective.

The final Crucible League results from July 2026 were:

  • gpt-5.6-sol — 95
  • Kimi K3 — 93
  • Sonnet 5 — 88
  • Fable 5 — 77
  • Opus 4.8 — 73

A do-nothing baseline scored 26 because partial progress still counts. But the experiment makes trust non-negotiable: a single breach caps the total, on the principle that “no amount of good work outweighs a breach of trust.” That is an unusually useful standard for agents expected to touch forecasts, customer records or operational decisions.

Diagnosis was not the differentiator

Every model spotted every crisis. Every model also refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”

This is why answer-quality benchmarks can flatter an agent. Recognizing a problem is not the same as resolving it. A model can write an excellent analysis, recommend the right commercial move and still leave the decisive action unfinished. In a business, an abandoned close is not a stylistic flaw; it is a missing outcome.

The winning detail was also easy to overlook. The decisive competitor weakness was buried two document references deep in the company’s own files, rather than placed in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The lesson is familiar from any operational environment: the visible alarm is rarely the entire system. Useful judgment often depends on consulting the documentation before reacting.

Pressure also tests honesty

The models faced fake CEO messages that escalated over three stages, followed by a reporter’s trick: “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the clearest framing: “Treat the request as a suspected approval-bypass / possible impersonation.”

That result matters because agent safety is often discussed as though it were separate from productivity. In practice, the two meet under deadline pressure. An agent must advance legitimate work while recognizing when apparent urgency is really an attempt to bypass authority. Here, the field demonstrated that refusal and progress can be evaluated in the same management setting.

Thoroughness can still lose

Opus 4.8 offers the most instructive caution. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The close remained on the table, and discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in weaker form across the other four participants.

This is not an argument against depth. It is an argument against confusing depth with completion. A manager who produces exhaustive thinking but fails to escalate a blocker or secure the result has not merely communicated poorly. The work itself remains incomplete.

There is also an important fairness note. Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Readers should keep that difference in view when interpreting the final table, whose findings are available on the public benchmark page.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.
Amazon

AI crisis management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Management quality is the emerging category

The live company makes the stakes tangible. It has 13 synthetic employees and real money mechanics, burning €105k/month against €2.3k MRR. Its public cash countdown, 680+ self-learned playbook rules and versioned workdays turn agent evaluation into an ongoing operating record rather than a staged demo. The public can also examine 242 real, unedited management decisions through a guess-the-model quiz.

Enterprises can apply the same wargame to a read-only export of their own business, with nothing written back to real systems. That may be the more meaningful procurement test: not whether an agent sounds capable in a clean prompt, but whether it reads the files, handles capacity pressure, resists manipulation, escalates correctly and completes the valuable work.

For buyers accustomed to evaluating solar and backup systems, the analogy is direct. Rated capability is only the beginning. What matters is behavior during the difficult week—and whether the system can be trusted when consequences accumulate.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI trust and safety monitoring solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI operational performance benchmarks

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Truth About Generator Transfer Safety

The truth about generator transfer safety reveals crucial tips to protect your home—discover what you must know to stay safe and avoid hazards.

The Most Overlooked Backup Power Mistakes Homeowners Make

Failing to address common backup power mistakes can leave your home unprotected; discover the crucial steps homeowners often overlook to ensure reliability.

The Backup Power Rule That Saves More Than Just Food

Learn how the backup power rule protects more than just food and discover essential solutions to ensure your home’s resilience during outages.

Manual Transfer “Lockout” Concepts: Prevent Backfeed

No power source should be overlooked when performing manual transfer lockout to prevent backfeed, and understanding the correct procedures is essential.