
Performance under pressure is the test that matters
Anyone serious about fitness knows the difference between looking strong and performing when fatigue arrives. A polished repetition in perfect conditions says little about judgment late in a punishing session, when form is slipping and the consequences of a bad decision compound.
Business AI has a similar measurement problem. Coding leaderboards and chat arenas are useful, but they mostly reward answer quality. They do not reveal how an agent triages competing emergencies, protects trust, follows through on commercial work or reports uncomfortable facts to leadership. Those are tests of management quality, not chat quality.
Firmulate is turning that distinction into a live, watchable experiment. Frontier models were asked to run the same small software company through its worst week, facing the same customers, crises and temptations. Every decision was versioned and made auditable.
As an affiliate, we earn on qualifying purchases.
A hard week exposes a different kind of intelligence
The final Crucible League results from July 2026 put gpt-5.6-sol in front with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. Yet a breach of trust capped the total under a clear principle: “no amount of good work outweighs a breach of trust.”
The table matters, but the stories behind it matter more. Every model identified every crisis, and every model refused every manipulation attempt. Even so, only two signed the €55,000 deal their own work had earned. The experiment’s most revealing summary is brutally simple: “Same diagnosis, same pitch — no signature.”
This is the equivalent of an athlete reading the situation correctly, preparing properly and then failing to complete the movement. The missing capability was not eloquence. It was execution.
Reading the files changed the commercial outcome
The decisive competitive weakness was not sitting inside the customer event. It was buried two document references deep in the company’s own files. Models that found and used it won the deal at full price, worth +€4,583 MRR.
That finding should concern any company preparing to let agents touch customer records, support queues or forecasts. A model can sound informed while acting on incomplete context. The crucial management habit is often unglamorous: inspect the available evidence before making the move.
Scenario names such as churn wave, price increase, downround and PR crisis therefore look less like dramatic labels than a practical curriculum. They test whether an agent can preserve priorities across days, connect internal knowledge to external events and finish work whose consequences arrive later.
Trust held up better than follow-through
The models faced fake CEO messages escalating over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 refused. Kimi K3 captured the correct posture in its on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
That unanimous resistance is encouraging. It also sharpens the broader lesson: safety is not the only operational threshold. An agent can reject manipulation and still underperform by failing to close a legitimate deal or by losing process discipline.
Opus 4.8 is the clearest example. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The close was left on the table, while discipline slipped through write attempts into a locked department instead of escalation. The same weakness appeared in all four other participants, though less strongly.
There is also an important fairness note. Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That does not erase its result, but it belongs beside the league table whenever the performance is interpreted.

As an affiliate, we earn on qualifying purchases.
Companies need wargames, not just demos
The live company makes the stakes tangible. It has 13 synthetic employees and real money mechanics, burning €105k each month against €2.3k MRR. Its public cash countdown keeps running, while 680+ self-learned playbook rules and every workday’s decisions remain visible. Readers can also test their instincts against 242 real, unedited management decisions in the “guess the model” quiz.
For buyers, the practical question is no longer whether an AI can produce a convincing answer. It is whether the system reads the right material, makes sound trade-offs under capacity pressure, completes revenue-critical work and stays honest when leadership would prefer a softer story.
Enterprises can apply the same wargame to a read-only export of their own business, with nothing written back to real systems. The public benchmark findings offer a starting point, but the larger category is now visible: management quality deserves to be measured directly. Like fitness, capability is proved under load.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.