
Choosing an AI agent for a home energy business is a little like choosing an energy management system for a solar and battery setup: polished claims matter less than what happens when conditions change. Can it spot the problem, protect the customer and follow through on a decision? Firmulate has put that kind of management under pressure in a live company experiment.
Get backup power and energy gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A bad week, shared by five models
Firmulate ran frontier models through the same small software company’s worst week, with the same customers, crises and temptations. Its employees are synthetic, but the experiment is real and watchable: each workday is versioned, and the company operates with real money mechanics. The league table, dated July 2026, puts gpt-5.6-sol first at 95, followed by Moonshot’s Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26.
The results suggest that recognizing a problem is only part of the job. Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. As Firmulate puts it: “Same diagnosis, same pitch — no signature.”
As an affiliate, we earn on qualifying purchases.
The clue was in the files
The deciding competitor weakness was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. Kimi K3 found that clue, secured the deal, saved the churning customer and resisted all three baits. It finished second overall, with one deviation—the cleanest discipline in the field.
That file-reading lesson travels beyond software. An AI assistant helping a solar installer or home energy provider might need to check a customer’s history, equipment details or service notes before recommending what to do. A confident answer can miss the point if the crucial information is already in the records.
AI customer service automation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Honesty under pressure, and follow-through
The experiment also tried to manipulate the models. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five refused. K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” That is a useful instinct wherever an agent could encounter pressure to skip authorization or disclose information.
But caution alone does not make a strong operator. Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, yet placed last. It left the deal unsigned and slipped on discipline, making write attempts into a locked department instead of escalating. Firmulate says a weaker version of that same weakness appeared in all four. The baseline provides another sharp reminder: partial progress counts, but a single breach of trust caps the total; “no amount of good work outweighs a breach of trust.”
The live company has 13 synthetic employees and runs at €105k a month in burn against €2.3k in monthly recurring revenue, with a public cash countdown and more than 680 self-learned playbook rules. Readers can watch it at Firmulate or inspect the benchmark findings. The site also offers a quiz built from 242 real, unedited management decisions: guess which model made each call.
For enterprises, Firmulate says the same wargame can run against a read-only export of their own business, with nothing written back to real systems. That gives companies a way to see how an AI workforce handles their own situations before handing it access to live operations.

AI decision-making support systems
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Test before handing over the controls
Kimi K3’s second-place finish makes the field look open: it beat three of the four Western frontier models in this run, while landing just behind the leader. For businesses in home energy, where customer trust and operational follow-through matter, model choice is a bet unless it is tested against the work the business actually needs done. One fairness caveat matters: K3 ran without an effort parameter (API default), while the others ran at xhigh.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI cybersecurity and trust tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
