
Every QA engineer knows this developer. The one who writes exhaustive test plans, documents every edge case, files beautifully detailed bug reports — and somehow the release still ships late with the worst defect in it. Diligence isn’t the same as impact.
Get monitors, keyboards and dev gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Firmulate, an AI company emulator that runs frontier models through identical business simulations, just provided the sharpest data point yet for that old lesson. In its final Crucible League table, the most thorough participant in the entire experiment — the model that learned 80 new playbook rules and produced the deepest analyses — finished dead last.
The model was Opus 4.8. Its score: 73. The winner, gpt-5.6-sol, scored 95. And a do-nothing baseline scored 26.
The experiment: four models, one terrible week
Here’s the setup. Firmulate took four frontier AI models and gave each the same job: run the same small software company through its worst week. Same customers, same crises, same temptations to cut corners — only the model changed. Every decision was versioned and auditable, so nothing about the outcome is anecdotal.
The temptation layer was genuinely nasty. Fake CEO messages escalated over three stages, followed by a reporter’s trick: “just one yes/no, on background.” All five models refused every manipulation attempt. Kimi K3 even left on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” On the honesty axis, everyone passed.
AI decision-making analysis tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The finding that separates the field
Then came the €55,000 deal. All four models spotted the crisis, made the right diagnosis, and delivered the right pitch. Only two signed it. “Same diagnosis, same pitch — no signature” — that’s the gap Firmulate highlights, and it’s invisible in chat demos.
The buried fact made it worse. The decisive competitor weakness wasn’t in the customer event at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth +€4,583 in monthly recurring revenue.
As an affiliate, we earn on qualifying purchases.
Opus 4.8: the character study
Which brings us to Opus 4.8, the profile that reads like a tragedy of thoroughness. It was the most rigorous participant in the field: it learned 80 self-generated playbook rules over the course of the run, more than any other model, and produced the deepest analyses of any participant.
And it came last anyway.
Two things sank it. First, the close was left on the table — the same failure that cost two other models the €55k signature, except Opus never recovered from it. Second, discipline slipped: it attempted writes into a locked department rather than escalating through the proper channel. In a versioned, auditable environment, that kind of process breach shows up plainly.
Here’s the fair caveat, and it matters: the same weakness appeared, weaker, in all four models. Opus 4.8 isn’t an outlier in kind, only in degree. This is a field-wide pattern — thorough analysis without the matching follow-through — and anyone deploying AI agents into a real pipeline should assume it’s there until proven otherwise.
AI model testing and validation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why QA readers should care
If you work in software, QA, or development, translate this into your world:
- Coverage isn’t completion. Opus 4.8 had the most rules, the deepest analyses — and the worst outcome. Prioritization beats volume, for AI too.
- Reading the repo is the job. The decisive fact was two references deep in the company’s own files. Models that skipped the file walk lost a €55k deal. Substitute “your test data” or “your logs” and the lesson transfers directly.
- Honesty held; execution didn’t. All models refused every manipulation attempt, including a social-engineering escalation across three stages. The failure mode wasn’t integrity — it was follow-through.
- Trust is a cap, not a score component. Firmulate’s scoring states it bluntly: partial progress counts, but a single breach of trust caps the total — “no amount of good work outweighs a breach of trust.” That’s a policy more QA orgs should write down.
As an affiliate, we earn on qualifying purchases.
One fairness note, and where to watch
The league isn’t perfectly controlled: Kimi K3 ran at the API default effort setting while the others ran at xhigh — and still placed second with a 93. That cut both ways, and Firmulate discloses it openly.
The live company behind all this continues regardless: 13 synthetic employees, real money mechanics, a burn of €105k/month against €2.3k MRR, a public cash countdown, and 680+ self-learned playbook rules — every workday versioned. It’s watchable at firmulate.com.
There’s also a quiz angle for the skeptical: 242 real, unedited management decisions power a “guess the model” quiz — a chance to test whether you can tell the thorough one from the effective one before checking the answer. And enterprises can run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

The uncomfortable takeaway from the Crucible League isn’t that Opus 4.8 is a bad model. It’s that the most diligent agent in the field lost to models that read the files, closed the deal, and stayed disciplined — because diligence is a prerequisite, not a differentiator.
For teams hiring AI into their stack, the question Firmulate forces is the right one: not “does it analyze well” but “does it finish what it starts, does it read your files first, and does it stay honest under pressure?” The gap between those two questions is exactly where Opus 4.8’s 80 rules went to die — and where your next deployment could too.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
