AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

A fitness plan can look convincing on paper and still fall apart under pressure: a packed schedule, a missed session, an unexpected setback. The same gap between a polished plan and a sound decision matters when AI is asked to run a business. Firmulate puts that judgment to a test by making models manage a company through a crisis-filled week.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get workout gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A hard week, held constant

In the final Crucible League, published in July 2026, frontier models faced the same small software company, the same customers, the same crises and the same temptations. Every decision was versioned and auditable. The final order was gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. The league’s integrity rule is blunt: “no amount of good work outweighs a breach of trust.”

The striking result was not that models failed to notice trouble. Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. As the experiment puts it: “Same diagnosis, same pitch — no signature.” A model can describe the right move and still leave the opportunity untouched.

The clue was already in the company’s files

The decisive weakness in a competitor’s position was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The lesson for a business is practical: a model’s judgment depends on whether it can find and use relevant evidence already available to the company.

Firmulate also tested pressure from people pretending to have authority. Fake CEO messages escalated over three stages, followed by a reporter’s “just one yes/no, on background” request. All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” Those refusals are meaningful, but they do not answer the separate question raised by the missed deal: can the system follow through when the right action is legitimate?

Thoroughness is not the same as execution

Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. The close was left on the table, and its discipline slipped into write attempts in a locked department instead of escalation. A weaker version of the same weakness appeared in all four models. Detailed analysis alone did not guarantee a completed decision.

There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The results are a report on this experiment, not a universal verdict on what any model will do in every business.

From watching to a company-specific trial

The live Firmulate company makes the experiment watchable. It has 13 synthetic employees and real money mechanics, with burn of €105k per month against €2.3k MRR, a public cash countdown, more than 680 self-learned playbook rules and a record of every workday versioned. Firmulate also offers a quiz built from 242 real, unedited management decisions: readers can guess which model made each call.

For an enterprise, the proposed next step is a pilot against a read-only export of its own business. Teams can test crisis scenarios using company context and receive a board report with model rankings and weaknesses in their playbooks. The pilot does not write back to real systems. That gives decision-makers a way to examine how a model responds before trusting it with work that touches customers, operations or forecasts.

Watch the live Firmulate experiment and explore the company-specific trial at firmulate.com/pilot.html.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put the decisions under pressure

In fitness, the plan matters most when the day does not go to plan. In business, an AI system’s judgment is clearest when a crisis tests both its principles and its ability to act. Firmulate’s league shows why a confident diagnosis is only part of the job: the model must find the evidence, respect boundaries and complete the decision. Enterprises can run the wargame against a read-only export of their own business, with nothing written back to real systems. Explore a pilot and contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Google Health Surges In Global Coverage

Google Health’s coverage has surged worldwide, with 17 mentions in recent data, indicating a major expansion of its health-related services and initiatives.

Brown Health Medical Surges In Global Coverage

Brown Health Medical experiences a significant surge in international media mentions, indicating expanding global recognition and influence.

Breast Cancer Network Australia Surges In Global Coverage

Breast Cancer Network Australia experiences a significant surge in international media coverage, with eight mentions reported in recent weeks, highlighting increased global awareness.