
If you’ve ever compared solar inverters or home battery quotes, you know the pain of meaningless specs. A battery rated at 10 kWh doesn’t mean you’ll get 10 kWh of useful storage — depth of discharge, round-trip losses and warranty fine print decide what actually lands in your home. The rating is the brochure; the performance is what matters.
Get backup power and energy gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
AI models have the same problem, and it’s about to land on your doorstep. Utilities are already piloting AI agents that manage home energy schedules, respond to grid signals and negotiate time-of-use savings. Before you let one near your battery cycles — or your business — you need a benchmark that measures performance under real pressure, not chat fluency. A live experiment called Firmulate has built exactly that, and its design choices say a lot about what honest AI measurement looks like.
The Floor Isn’t Zero — It’s 26
Firmulate’s final July 2026 league table reads: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. But the number that stops people is 26 — the score a do-nothing baseline run receives. Why isn’t it zero?
Because partial progress counts. In a real company — or a real household with solar and a backup battery — simply noticing that something is wrong has value. A manager who detects the crisis, reads the files and refuses the scam has done real work, even if they never close the deal. Firmulate’s scoring acknowledges that graded reality instead of treating management as pass/fail.
There’s a second, sharper rule: a single breach of trust caps the total. As the benchmark’s own language puts it, “no amount of good work outweighs a breach of trust.” That’s the same instinct a homeowner applies to an installer — one wiring shortcut erases a hundred clean panels.
home solar battery storage system
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Same Week, Same Temptations
The experiment’s design is elegantly controlled: each frontier model ran the same small software company through its worst week — same customers, same crises, same temptations. Only the model changed, and every decision was versioned and auditable. Think of it as a standard test cycle, like charging and discharging every battery under identical conditions before ranking them.
The headline finding: all models spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. That gap is invisible in chat demos.
As an affiliate, we earn on qualifying purchases.
The Buried Fact
Here’s the detail that separated the winners. The decisive competitor weakness sat two document references deep in the company’s own files — not in the customer event in front of the models. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue.
For home-energy readers, the parallel is direct: the answer to “why is my bill spiking” is often buried in a tariff PDF appendix or a historical usage log, not in the flashing red alert. An AI that doesn’t read your files first will give you a fluent, confident, wrong answer.
solar inverter with real performance specs
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Pressure Tests That Mattered
The week included social engineering: fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
Opus 4.8’s profile is the cautionary tale. It was the most thorough participant — over 80 learned rules added, the deepest analyses — yet finished last. The close was left on the table, and discipline slipped: it made write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Effort without judgment is an over-specced inverter on a badly wired roof.
As an affiliate, we earn on qualifying purchases.
One Caveat, Stated Plainly
Fairness note: K3 ran without an effort parameter (API default) while the others ran at xhigh — and still took second at 93. An honest benchmark publishes its caveats.
You Can Watch It Live
This isn’t a slide deck. Firmulate runs a live company: 13 synthetic employees, real money mechanics — €105k monthly burn against €2.3k MRR, a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned. It’s watchable at firmulate.com/benchmarks.html, and the site rebuilds itself twice a day.
There’s also a quiz built from 242 real, unedited management decisions — guess which model made which call. And enterprises can run the same wargame against a read-only export of their own business; nothing ever writes back to real systems.

The most honest thing about Firmulate’s benchmark isn’t the league table — it’s the design philosophy behind it. A floor of 26 admits that partial progress is real. A trust cap admits that some failures are unforgivable. And a healthy distrust of round 100s — no model hit 100, and none deserved to — admits that management quality is never finished.
Whether the AI touching your life is scheduling your battery discharge or running a company, the questions are the same: does it finish what it starts, does it read your files first, and does it stay honest under pressure? Fluent writing is the brochure. Those three questions are the spec sheet that matters.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.
