firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get monitors, keyboards and dev gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A Benchmark That Doesn’t Start at Zero

If you’ve ever run a QA suite, you know the temptation of a perfect score. Everything passes, the dashboard glows green, and someone asks whether the tests were too easy. The team behind Firmulate’s benchmark league built their scoring with the same skepticism — including a built-in distrust of round 100s. Their most interesting design decision is at the other end of the scale, though: an AI manager that does absolutely nothing still scores 26 points. Here’s why that’s a feature, not a bug.

Amazon

AI management benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Do-Nothing Baseline

Firmulate ran an experiment where each frontier AI model was given the same job: run the same small software company through its worst week. Same customers, same crises, same temptations to cheat — only the model changes. Every decision was versioned and auditable, so nothing about a run is hand-waved.

Before scoring any model, the team established a floor: what happens if the manager does nothing at all? The answer is 26 points. That number exists because partial progress counts. A company that limps through a crisis week with an inert manager isn’t at zero — some things get half-handled, some fires burn out on their own. A benchmark that awards 0 to a do-nothing run pretends the world is binary, and management never is.

The ceiling has its own logic. A single breach of trust caps the total grade — in the team’s words, “no amount of good work outweighs a breach of trust.” So the scale runs from “inert but not destructive” to “excellent, provided nothing disqualifying happened.” It’s the kind of scoring rubric you’d want applied to a human manager, too.

Amazon

AI decision-making analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Final League, July 2026

The Crucible League finished with gpt-5.6-sol first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. Note that no one hit 100 — and the benchmark design wouldn’t particularly trust one if it appeared.

The headline finding cuts to the heart of what these systems can and can’t do. All four models spotted every crisis and refused every manipulation attempt. Only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. That gap is invisible in chat demos.

Amazon

AI documentation reading tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Buried Fact

The decisive detail in the deal sat two document references deep in the company’s own files — not in the customer event. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. For anyone who’s watched an integration fail because nobody read the existing documentation, this will feel familiar: the work was never in the conversation. It was in the repository.

Amazon

AI performance evaluation platforms

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Thoroughness Isn’t Enough

Opus 4.8 is the cautionary tale of the league. It was the most thorough participant — over 80 learned rules and the deepest analyses — yet finished last. The close was left on the table, and discipline slipped: it made write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Reading everything and deciding nothing is a failure mode QA teams know well.

One fairness note the team disclosed: Kimi K3 ran without an effort parameter (API default) while the others ran at xhigh — and still placed second.

Under Social Pressure

The week included fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” Five of five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That reasoning is exactly what you’d want logged in an audit trail for any AI touching production systems.

See It Yourself

The experiment didn’t end with the league. Firmulate runs a live synthetic company — 13 employees, real money mechanics, burning €105k a month against €2.3k MRR, with a public cash countdown and 680+ self-learned playbook rules, every workday versioned. It’s watchable at firmulate.com/live. There’s also a “guess the model” quiz powered by 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

An honest benchmark tells you three things a chat demo can’t: where the floor is, what partial progress is worth, and what a single breach of trust costs. Firmulate’s 26-point do-nothing baseline is a quiet admission that evaluation is hard — and that a suspicious eye toward perfect scores is a sign the graders are paying attention. If you’re evaluating AI agents for work that touches real systems, that’s the standard to hold them to.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Reddit Imposes New Restrictions On Old Reddit – Engadget

Reddit plans to limit Old Reddit access to people who used it in the past six months and end RSS support on November 13, citing scraping concerns.

DARPA Funds Qunnect To Improve Quantum Network Reliability

DARPA has allocated funding to Qunnect to develop technologies aimed at improving the stability and reliability of quantum communication networks.

Best Automated Testing Tools Compared

Compare Playwright and Selenium on browser coverage, reliability, setup, ecosystem, and cost to choose the right automated testing tool.

The Work By Valve’s Timur Kristóf On Improving Old AMD GPUs On Linux

Timur Kristóf presented work on AMDGPU kernel support for decade-old GCN 1.0 and 1.1 cards, including display, power-management and reset fixes.