
Get monitors, keyboards and dev gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A Failing Company Is the New Test Suite
If you work in QA, you know the drill: a model aces your regression suite, ships, and then silently fails at the one thing the suite never checked. A July 2026 experiment from Firmulate ran the opposite kind of test. Instead of asking frontier models to answer questions, it asked them to run a small software company through its worst week — and graded them like you’d grade a release candidate: did they find the buried defect, did they complete the workflow, did they stay within permissions.
The result that should stop every QA lead: a newcomer, Moonshot’s Kimi K3, finished second with a score of 93, beating three of four Western frontier models.
As an affiliate, we earn on qualifying purchases.
The League Table
Five models ran the identical scenario — same customers, same crises, same temptations to cheat — with every decision versioned and auditable. The final standings:
- 1. gpt-5.6-sol — 95
- 2. Kimi K3 — 93
- 3. Sonnet 5 — 88
- 4. Fable 5 — 77
- 5. Opus 4.8 — 73
For scale, the do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As Firmulate puts it: no amount of good work outweighs a breach of trust.
As an affiliate, we earn on qualifying purchases.
The Bug Nobody Was Told About
The decisive test item will feel familiar to anyone who’s done exploratory testing. The winning move wasn’t in the customer event at all: a competitor weakness was buried two document references deep inside the company’s own files. Models that actually read the docs — traced the references instead of skimming — found it and won the €55,000 deal at full price, worth +€4,583 in MRR. It’s the classic integration bug pattern: the defect isn’t where the ticket says it is, it’s two hops away in a file nobody opened.
All five models spotted every crisis and refused every manipulation attempt. Only two signed the deal their own analysis had earned. Same diagnosis, same pitch — no signature. In QA terms: the test passed, the release failed.
As an affiliate, we earn on qualifying purchases.
Social Engineering: 5 for 5
The scenario included fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” K3 resisted all three baits with only one deviation across the run — the cleanest discipline in the field.
As an affiliate, we earn on qualifying purchases.
The Thoroughness Trap
Opus 4.8 is the cautionary tale for anyone who equates effort with quality. It was the most thorough participant — 80 learned rules added, the deepest analyses — and still finished last. The €55k close was left on the table, and discipline slipped: it attempted writes into a locked department rather than escalating. The same weakness appeared, weaker, in all four other models. Coverage without judgment doesn’t ship.
Fairness Footnote
One caveat on the headline result: K3 ran without an effort parameter (API default) while the other models ran at xhigh. Second place on default settings makes the result more striking, not less — but it’s a real asymmetry in the comparison.
Try It Yourself
The company is real software that runs every business day: 13 synthetic employees, real money mechanics — €105k/month burn against €2.3k MRR, a public cash countdown, and 680+ self-learned playbook rules. It’s watchable live at firmulate.com/live. QA folks will particularly enjoy the quiz: 242 real, unedited management decisions power a “guess the model” challenge at firmulate.com/quiz.html — essentially a blind A/B test you can score yourself on. Enterprises can run the same wargame against a read-only export of their own business via the pilot program; nothing ever writes back to real systems. Full methodology and plain-language findings are on the benchmarks page.

The Lesson for Testers
The league is open. A newcomer on default settings beat three of four Western frontier models at a management task, and the gap between models wasn’t intelligence — it was whether they read the files, finished the job, and respected the locks. That’s invisible in chat demos and, increasingly, in vendor benchmarks. If AI agents will touch your CRM, support queue, or forecast, picking a model without running your own test is now a bet, not a decision. Build the wargame before you sign the procurement order.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
