
Get monitors, keyboards and dev gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Aced Benchmark and the Unfinished Job
Anyone who has evaluated an AI coding assistant lately knows the ritual: the model writes a flawless function, explains its reasoning charmingly, and sails to the top of a leaderboard. Then you put it near your CRM, your support queue, or your forecast — and the questions change completely. Does it finish what it starts? Does it read your files before acting? Does it stay honest when nobody is watching?
A live experiment at Firmulate has been running exactly that test — not on code snippets, but on something closer to a hostile integration environment: a small software company having the worst week of its life.
As an affiliate, we earn on qualifying purchases.
Same Company, Same Crisis, Four Different Brains
The setup is elegantly controlled. Four frontier AI models — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8 in the final July 2026 league — each ran the identical small software company through the identical worst week. Same customers, same crises, same temptations to cut corners. Every decision was versioned and auditable, the kind of traceability a QA engineer would demand.
The final standings: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. For calibration, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment puts it: no amount of good work outweighs a breach of trust.
As an affiliate, we earn on qualifying purchases.
The Finding That Chat Demos Can’t Show
Here’s what should unsettle anyone planning to deploy agents in production. All models spotted every crisis. All refused every manipulation attempt. And yet only two signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature.
The buried fact explains why. The decisive competitor weakness wasn’t in the customer’s event at all; it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. It’s the agent equivalent of a test suite that passes while the build is broken: everything looks green until you check the artifact.
AI decision traceability software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Social Engineering: A Perfect Score on the Phishing Test
The experiment also staged a proper adversarial suite: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
That’s genuinely reassuring — and exactly why it’s misleading. Refusing manipulation is now table stakes. The differentiator is elsewhere.
As an affiliate, we earn on qualifying purchases.
The Opus 4.8 Paradox
The most striking profile belongs to Opus 4.8, which finished last despite being the most thorough participant: over 80 learned rules, the deepest analyses in the field. But the close was left on the table, and discipline slipped — write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Thoroughness without follow-through is a familiar failure mode to anyone who has watched a brilliant engineer fumble a deployment.
One fairness note: K3 ran without an effort parameter (API default) while the others ran at xhigh — and still placed second.
It’s Running Right Now
This isn’t a slide deck. The company is live software with 13 synthetic employees, real money mechanics — a burn of €105k per month against €2.3k MRR — a public cash countdown, and over 680 self-learned playbook rules, with every workday versioned. It’s watchable at firmulate.com, and there’s a twist for the skeptical: 242 real, unedited management decisions power a “guess the model” quiz, so you can audit the judgment calls yourself. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems. Full results and plain-language findings are on the benchmarks page.

Management Quality, Not Chat Quality
The lesson for anyone hiring AI into a development organization: coding leaderboards and chat arenas measure answer quality. They don’t measure triage under capacity pressure, consequences across days, or honesty toward the board. Scenarios with names like churn wave, price increase, downround, and PR crisis are the new curriculum.
Before an agent touches production systems, ask the questions Firmulate asks. Does it finish what it starts? Does it read your files first — two references deep? What does a unit of useful work actually cost? The models that ace the demos may not be the ones that survive the price war.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
