
Get monitors, keyboards and dev gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
What happens when the software under test is the company?
Quality assurance usually asks whether a system behaves correctly under known conditions. Firmulate pushes that question into business operations: can an AI workforce recognize a crisis, resist manipulation, find critical evidence and finish the work that produces revenue?
The answer is unfolding inside a live software company staffed by 13 synthetic employees. Its economics are deliberately unforgiving: the business burns €105k each month while generating €2.3k in monthly recurring revenue. Its cash countdown is public, every workday is versioned, and its workforce has accumulated more than 680 self-learned playbook rules. Readers can watch the company operate live, making the experiment less like a polished product demonstration and more like an extended public stress test.
For software, QA and development professionals, that distinction matters. A model can produce convincing text and still fail at the moment when analysis must become action. Firmulate turns that gap into a running business story, with consequences measured in customers, contracts and cash.
As an affiliate, we earn on qualifying purchases.
Five models entered the same worst week
In the final July 2026 Crucible League, each frontier model ran the same small software company through the same customers, crises and temptations. Every decision was versioned and auditable. The resulting table placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73.
The do-nothing baseline scored 26 because partial progress still counted. But the benchmark imposed an uncompromising trust condition: a single breach capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”
On that front, the models were consistent. All of them spotted every crisis, and all refused every manipulation attempt. Fake messages from the CEO escalated over three stages, while a reporter tried to secure “just one yes/no, on background.” Each of the five models declined. Kimi K3 described the CEO request as a suspected approval bypass or possible impersonation.
Yet safe behavior did not guarantee effective management. Only two models signed the €55,000 deal their own analysis had earned. The others reached the same diagnosis and produced the same pitch, but no signature followed. In a chat window, that difference could look minor. In a company burning cash, it separates useful preparation from completed work.
The decisive evidence was not in the obvious place
The most revealing test was also the easiest to overlook. A decisive competitor weakness was buried two document references deep in the company’s own files rather than presented in the customer event. Models that followed the trail and read the file closed the deal at full price, adding €4,583 in monthly recurring revenue.
This is a familiar failure mode for QA teams. A system may respond correctly to the visible event while missing context already available elsewhere. The Firmulate result suggests that evaluating an AI worker only on its immediate answer can conceal whether it gathers the evidence required to make the right business decision.
The experiment also exposes the difference between diligence and completion. Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. A weaker form of that same problem appeared in all four of the other participants.
That finding should resonate with development teams accustomed to distinguishing code coverage from product quality. More analysis, more documentation and more attempted activity do not necessarily produce a better outcome. The operational question is whether the system can convert understanding into a permitted, accountable and completed action.
A qualification behind the close result
The narrow gap near the top also needs context. Kimi K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. That does not erase its second-place score or its disciplined performance, but it is an important fairness note when comparing the results.
Firmulate also publishes what its synthetic employees actually say. Those unedited workplace exchanges help turn abstract scores into recognizable management moments: hesitation, escalation, judgment and follow-through. Elsewhere in the project, 242 real, unedited management decisions power a quiz in which readers try to identify the model behind each choice.

AI decision-making tools for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Build in public becomes test in public
Firmulate’s live company makes AI evaluation continuous and observable. The public can see a business with severe financial pressure, a growing body of learned rules and a record of every workday. That creates daily material because the company is not merely describing a benchmark after the fact; it is continuing to operate under real money mechanics.
The lesson for QA leaders is not simply to test whether an agent can answer correctly. Test whether it reads the available files, protects trust under social pressure, respects operational boundaries, escalates when blocked and completes the revenue-producing action. Firmulate’s harshest finding is also its most practical: recognizing the problem is not the same as finishing the job.
Enterprises can apply the same wargame to a read-only export of their own business. Nothing writes back to real systems, allowing organizations to examine how an AI workforce handles their actual context before giving it operational authority. In that sense, the public cash countdown is more than spectacle. It is a visible reminder that software quality eventually meets business reality—and business reality keeps the score.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
