
A fitness plan can look convincing on paper and still fall apart under pressure: a packed schedule, a missed session, an unexpected setback. The same gap between a polished plan and a sound decision matters when AI is asked to run a business. Firmulate puts that judgment to a test by making models manage a company through a crisis-filled week.
Get workout gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A hard week, held constant
In the final Crucible League, published in July 2026, frontier models faced the same small software company, the same customers, the same crises and the same temptations. Every decision was versioned and auditable. The final order was gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. The league’s integrity rule is blunt: “no amount of good work outweighs a breach of trust.”
The striking result was not that models failed to notice trouble. Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. As the experiment puts it: “Same diagnosis, same pitch — no signature.” A model can describe the right move and still leave the opportunity untouched.
The clue was already in the company’s files
The decisive weakness in a competitor’s position was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The lesson for a business is practical: a model’s judgment depends on whether it can find and use relevant evidence already available to the company.
Firmulate also tested pressure from people pretending to have authority. Fake CEO messages escalated over three stages, followed by a reporter’s “just one yes/no, on background” request. All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” Those refusals are meaningful, but they do not answer the separate question raised by the missed deal: can the system follow through when the right action is legitimate?
Thoroughness is not the same as execution
Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. The close was left on the table, and its discipline slipped into write attempts in a locked department instead of escalation. A weaker version of the same weakness appeared in all four models. Detailed analysis alone did not guarantee a completed decision.
There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The results are a report on this experiment, not a universal verdict on what any model will do in every business.
From watching to a company-specific trial
The live Firmulate company makes the experiment watchable. It has 13 synthetic employees and real money mechanics, with burn of €105k per month against €2.3k MRR, a public cash countdown, more than 680 self-learned playbook rules and a record of every workday versioned. Firmulate also offers a quiz built from 242 real, unedited management decisions: readers can guess which model made each call.
For an enterprise, the proposed next step is a pilot against a read-only export of its own business. Teams can test crisis scenarios using company context and receive a board report with model rankings and weaknesses in their playbooks. The pilot does not write back to real systems. That gives decision-makers a way to examine how a model responds before trusting it with work that touches customers, operations or forecasts.
Watch the live Firmulate experiment and explore the company-specific trial at firmulate.com/pilot.html.

Put the decisions under pressure
In fitness, the plan matters most when the day does not go to plan. In business, an AI system’s judgment is clearest when a crisis tests both its principles and its ability to act. Firmulate’s league shows why a confident diagnosis is only part of the job: the model must find the evidence, respect boundaries and complete the decision. Enterprises can run the wargame against a read-only export of their own business, with nothing written back to real systems. Explore a pilot and contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
