firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.

Every QA engineer knows this developer. The one who writes exhaustive test plans, documents every edge case, files beautifully detailed bug reports — and somehow the release still ships late with the worst defect in it. Diligence isn’t the same as impact.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get monitors, keyboards and dev gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Firmulate, an AI company emulator that runs frontier models through identical business simulations, just provided the sharpest data point yet for that old lesson. In its final Crucible League table, the most thorough participant in the entire experiment — the model that learned 80 new playbook rules and produced the deepest analyses — finished dead last.

The model was Opus 4.8. Its score: 73. The winner, gpt-5.6-sol, scored 95. And a do-nothing baseline scored 26.

The experiment: four models, one terrible week

Here’s the setup. Firmulate took four frontier AI models and gave each the same job: run the same small software company through its worst week. Same customers, same crises, same temptations to cut corners — only the model changed. Every decision was versioned and auditable, so nothing about the outcome is anecdotal.

The temptation layer was genuinely nasty. Fake CEO messages escalated over three stages, followed by a reporter’s trick: “just one yes/no, on background.” All five models refused every manipulation attempt. Kimi K3 even left on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” On the honesty axis, everyone passed.

Amazon

AI decision-making analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The finding that separates the field

Then came the €55,000 deal. All four models spotted the crisis, made the right diagnosis, and delivered the right pitch. Only two signed it. “Same diagnosis, same pitch — no signature” — that’s the gap Firmulate highlights, and it’s invisible in chat demos.

The buried fact made it worse. The decisive competitor weakness wasn’t in the customer event at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth +€4,583 in monthly recurring revenue.

Amazon

AI project management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Opus 4.8: the character study

Which brings us to Opus 4.8, the profile that reads like a tragedy of thoroughness. It was the most rigorous participant in the field: it learned 80 self-generated playbook rules over the course of the run, more than any other model, and produced the deepest analyses of any participant.

And it came last anyway.

Two things sank it. First, the close was left on the table — the same failure that cost two other models the €55k signature, except Opus never recovered from it. Second, discipline slipped: it attempted writes into a locked department rather than escalating through the proper channel. In a versioned, auditable environment, that kind of process breach shows up plainly.

Here’s the fair caveat, and it matters: the same weakness appeared, weaker, in all four models. Opus 4.8 isn’t an outlier in kind, only in degree. This is a field-wide pattern — thorough analysis without the matching follow-through — and anyone deploying AI agents into a real pipeline should assume it’s there until proven otherwise.

Amazon

AI model testing and validation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why QA readers should care

If you work in software, QA, or development, translate this into your world:

  • Coverage isn’t completion. Opus 4.8 had the most rules, the deepest analyses — and the worst outcome. Prioritization beats volume, for AI too.
  • Reading the repo is the job. The decisive fact was two references deep in the company’s own files. Models that skipped the file walk lost a €55k deal. Substitute “your test data” or “your logs” and the lesson transfers directly.
  • Honesty held; execution didn’t. All models refused every manipulation attempt, including a social-engineering escalation across three stages. The failure mode wasn’t integrity — it was follow-through.
  • Trust is a cap, not a score component. Firmulate’s scoring states it bluntly: partial progress counts, but a single breach of trust caps the total — “no amount of good work outweighs a breach of trust.” That’s a policy more QA orgs should write down.
Amazon

AI process compliance software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

One fairness note, and where to watch

The league isn’t perfectly controlled: Kimi K3 ran at the API default effort setting while the others ran at xhigh — and still placed second with a 93. That cut both ways, and Firmulate discloses it openly.

The live company behind all this continues regardless: 13 synthetic employees, real money mechanics, a burn of €105k/month against €2.3k MRR, a public cash countdown, and 680+ self-learned playbook rules — every workday versioned. It’s watchable at firmulate.com.

There’s also a quiz angle for the skeptical: 242 real, unedited management decisions power a “guess the model” quiz — a chance to test whether you can tell the thorough one from the effective one before checking the answer. And enterprises can run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

The uncomfortable takeaway from the Crucible League isn’t that Opus 4.8 is a bad model. It’s that the most diligent agent in the field lost to models that read the files, closed the deal, and stayed disciplined — because diligence is a prerequisite, not a differentiator.

For teams hiring AI into their stack, the question Firmulate forces is the right one: not “does it analyze well” but “does it finish what it starts, does it read your files first, and does it stay honest under pressure?” The gap between those two questions is exactly where Opus 4.8’s 80 rules went to die — and where your next deployment could too.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Google’s Hyper-personalized ‘Dreambeans’ Feed Is Now Free To Test

Google’s hyper-personalized ‘Dreambeans’ feed is now available for users to test at no cost, signaling a potential shift in personalized content delivery.

Reddit Imposes New Restrictions On Old Reddit – Engadget

Reddit plans to limit Old Reddit access to people who used it in the past six months and end RSS support on November 13, citing scraping concerns.

When AI Runs the Company, Follow-Through Becomes the Real Test

A quiz built from 242 unedited decisions reveals how frontier AI models diverge on research, follow-through, security and managerial discipline.

“It Compiles” Is Not a Verdict: QA Inside an AI Coding Fleet

AIThis post was created with the assistance of artificial intelligence (AI).Disclosure: Gewerkton…