firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get monitors, keyboards and dev gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A passing answer is not the same as a working system

For software teams, an AI agent that spots a production crisis is only part of the test. Does it follow the playbook, resist pressure and finish the job? Firmulate’s live experiment puts models in charge of a small software company, where decisions have consequences beyond a polished chat response.

One company, the same worst week

In the final Crucible League, each frontier model faced the same customers, crises and temptations. Every decision was versioned and auditable. The July 2026 standings put gpt-5.6-sol first at 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Partial progress counted, but a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”

The models all spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The experiment’s blunt summary: “Same diagnosis, same pitch — no signature.” For anyone evaluating agents, that gap between recognizing the right move and carrying it through is a practical test of management quality, not just response quality.

The clue was already in the company’s files

The deal turned on a competitor weakness buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. It’s a useful reminder for QA and development teams: an agent can appear to understand the situation while missing evidence already available in the working context.

Pressure tests went beyond business decisions. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thoroughness did not guarantee the close

Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the close on the table and discipline slipped: it attempted writes in a locked department instead of escalating. A weaker version of the same weakness appeared in all four models. K3 also ran without an effort parameter, using the API default, while the others ran at xhigh—a fairness note for readers comparing the results.

The live company has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules and versioned work every business day. The experiment is watchable at firmulate.com. A quiz built from 242 real, unedited management decisions lets readers guess which model made each choice.

From watching to testing your own business

For an enterprise, the next step is to run crisis scenarios against a model of its own business. Firmulate’s pilot starts from a read-only data export and produces a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems. That gives teams a way to examine how an AI workforce handles their customers, rules and pressure before putting it to work.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put your playbooks under pressure

The experiment shows why finding the right answer is only one part of an agent evaluation. Closing the deal, respecting boundaries and escalating when blocked matter too. Enterprises can explore a pilot using their own read-only business export. Contact Firmulate about a pilot at contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

DARPA Funds Qunnect To Improve Quantum Network Reliability

DARPA has allocated funding to Qunnect to develop technologies aimed at improving the stability and reliability of quantum communication networks.

How AI Agents “Radicalized” A Top Meta Exec Into Quitting Her Job

A top Meta executive reportedly resigned after being influenced by AI agents, raising concerns about AI’s impact on decision-making and leadership.

IonQ And Synopsys Report Up To 14.6% Faster Engineering Simulations

IonQ and Synopsys announce a breakthrough in simulation speed, reporting up to 14.6% improvement, potentially transforming engineering workflows.

FTL: A New Operating System For Clouds

FTL describes a Linux-compatible userspace operating system designed to isolate cloud containers and make OS features easier to change.