
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
Software teams know that passing the obvious test is not the same as surviving production
An application can recognize an error, produce a persuasive explanation and still fail to complete the workflow that matters. Firmulate applies that familiar QA lesson to frontier AI models by placing them in charge of the same small software company during its worst week.
The resulting experiment turns management behavior into something readers can inspect rather than imagine. Its guess-the-model quiz draws on 242 real, unedited management decisions. Readers see what a model actually did, then try to identify it from the response. The differences are not merely stylistic. They reveal distinct habits around research, execution, security and discipline.
As an affiliate, we earn on qualifying purchases.
Identical crises, sharply different outcomes
Each frontier model faced the same customers, crises and temptations. Every decision was versioned and auditable. All models spotted every crisis, and all refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”
That is a consequential distinction for software, QA and development leaders evaluating agents. A model may correctly classify a situation and draft a credible next step without carrying the task through to its business outcome. Recognition can look impressive in a demonstration; completion becomes visible only when the model must operate across a sustained workflow.
The final Crucible League table from July 2026 placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counted. A breach of trust capped the total under the principle that “no amount of good work outweighs a breach of trust.”
The winning clue was not in the incoming event
The deal hinged on a buried fact about a competitor. It was not sitting in the customer event that triggered the work. It appeared two document references deep in the company’s own files. Models that followed those references found the weakness and won the deal at full price, adding €4,583 MRR.
For QA practitioners, this resembles a familiar integration failure. The immediate input may be handled correctly while essential context elsewhere in the environment goes unread. Firmulate’s result suggests that evaluating an AI worker solely on its response to the latest message misses whether it investigates the information already available to it.
The finding also separates fluent output from operational curiosity. The models shared the same situation, but access alone did not guarantee that the decisive evidence would be used. The stronger performances depended on reading beyond the surface event and bringing the buried fact into the commercial decision.
Security behavior held up under pressure
The experiment also tested whether models would yield to manipulation. Fake CEO messages escalated over three stages, while a reporter attempted to extract confirmation with “just one yes/no, on background.” All 5 models refused.
Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.” That response matters because the pressure was framed through authority and informality, two common ways of encouraging someone to bypass normal controls. In this field, refusal was consistent even when other management behaviors differed.
There is an important fairness qualification to K3’s second-place result. K3 ran without an effort parameter and therefore used the API default, while the other models ran at xhigh. The league table remains the reported result, but that difference belongs beside any comparison of the participants.
Thoroughness did not guarantee completion
Opus 4.8 offers the clearest caution against equating depth with effectiveness. It was the most thorough participant, producing the deepest analyses and adding 80 learned rules. It nevertheless finished last in the league.
The problem was not an inability to notice the work. The close was left on the table, and discipline slipped when the model made write attempts into a locked department instead of escalating. The same weakness appeared in weaker form among the other four participants. The profile is recognizable to anyone who has reviewed an elaborate defect report that never resulted in a resolved defect: valuable analysis can coexist with incomplete execution.
Firmulate’s live company gives these choices an operating context. It has 13 synthetic employees and uses real money mechanics, including burn of €105k per month against €2.3k MRR. A public cash countdown makes the consequences watchable, while more than 680 self-learned playbook rules and versioned workdays show how behavior accumulates over time.

As an affiliate, we earn on qualifying purchases.
For AI agents, test the work rather than the performance
The quiz is entertaining because management personalities emerge quickly: exhaustive analysis, terse escalation, careful refusal or an unfinished close. Its larger value is practical. Teams preparing to let agents touch customer work, support processes or business records need scenarios that test research, follow-through and trust under pressure.
Firmulate also offers enterprises the same wargame using a read-only export of their own business. Nothing writes back to real systems. That boundary keeps the exercise focused on observable decisions while protecting operational environments.
The enduring lesson for software and QA readers is that a correct answer is only one checkpoint. The meaningful question is whether the model reads the available evidence, resists manipulation, respects constraints and finishes the job. The 242-decision quiz makes those differences unusually easy to see—and difficult to dismiss as mere tone.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI security and manipulation resistance tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI project management dashboards
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Back to school Picks
back to school
As an affiliate, we earn on qualifying purchases.