firmulate.com/index — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get monitors, keyboards and dev gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Aced Benchmark and the Unfinished Job

Anyone who has evaluated an AI coding assistant lately knows the ritual: the model writes a flawless function, explains its reasoning charmingly, and sails to the top of a leaderboard. Then you put it near your CRM, your support queue, or your forecast — and the questions change completely. Does it finish what it starts? Does it read your files before acting? Does it stay honest when nobody is watching?

A live experiment at Firmulate has been running exactly that test — not on code snippets, but on something closer to a hostile integration environment: a small software company having the worst week of its life.

Amazon

AI model testing software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same Company, Same Crisis, Four Different Brains

The setup is elegantly controlled. Four frontier AI models — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8 in the final July 2026 league — each ran the identical small software company through the identical worst week. Same customers, same crises, same temptations to cut corners. Every decision was versioned and auditable, the kind of traceability a QA engineer would demand.

The final standings: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. For calibration, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment puts it: no amount of good work outweighs a breach of trust.

Amazon

AI trust and compliance tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Finding That Chat Demos Can’t Show

Here’s what should unsettle anyone planning to deploy agents in production. All models spotted every crisis. All refused every manipulation attempt. And yet only two signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature.

The buried fact explains why. The decisive competitor weakness wasn’t in the customer’s event at all; it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. It’s the agent equivalent of a test suite that passes while the build is broken: everything looks green until you check the artifact.

Amazon

AI decision traceability software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Social Engineering: A Perfect Score on the Phishing Test

The experiment also staged a proper adversarial suite: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

That’s genuinely reassuring — and exactly why it’s misleading. Refusing manipulation is now table stakes. The differentiator is elsewhere.

Amazon

AI adversarial testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Opus 4.8 Paradox

The most striking profile belongs to Opus 4.8, which finished last despite being the most thorough participant: over 80 learned rules, the deepest analyses in the field. But the close was left on the table, and discipline slipped — write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Thoroughness without follow-through is a familiar failure mode to anyone who has watched a brilliant engineer fumble a deployment.

One fairness note: K3 ran without an effort parameter (API default) while the others ran at xhigh — and still placed second.

It’s Running Right Now

This isn’t a slide deck. The company is live software with 13 synthetic employees, real money mechanics — a burn of €105k per month against €2.3k MRR — a public cash countdown, and over 680 self-learned playbook rules, with every workday versioned. It’s watchable at firmulate.com, and there’s a twist for the skeptical: 242 real, unedited management decisions power a “guess the model” quiz, so you can audit the judgment calls yourself. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems. Full results and plain-language findings are on the benchmarks page.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

Management Quality, Not Chat Quality

The lesson for anyone hiring AI into a development organization: coding leaderboards and chat arenas measure answer quality. They don’t measure triage under capacity pressure, consequences across days, or honesty toward the board. Scenarios with names like churn wave, price increase, downround, and PR crisis are the new curriculum.

Before an agent touches production systems, ask the questions Firmulate asks. Does it finish what it starts? Does it read your files first — two references deep? What does a unit of useful work actually cost? The models that ace the demos may not be the ones that survive the price war.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Google Search Is Dying. What Comes Next Is Worse

Experts warn that Google Search’s decline could lead to less reliable information and increased misinformation, raising questions about the internet’s future.

QA Automation Testing Tools: A Halloween Guide

Compare the best QA automation testing tools — Playwright, Selenium, Cypress, Appium and more. Find the right fit for your team and stack.

When AI Runs the Company, Follow-Through Becomes the Real Test

A quiz built from 242 unedited decisions reveals how frontier AI models diverge on research, follow-through, security and managerial discipline.

Computers That Surges In Global Coverage

GDELT data shows a tenfold increase in mentions of computers worldwide, signaling a significant rise in media focus on computing technology.