
Just like in fitness, where sheer effort doesn’t guarantee progress, in AI-driven decision-making, diligence alone isn’t enough to guarantee success. Recent experiments reveal that even the most detailed AI models can fall short when it counts — showing us that focus and prioritization matter just as much as effort.
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
The Firmulate Experiment: Putting AI to the Test in a Small Software Business
Imagine a real-world scenario where AI models manage a small software company’s week—dealing with customer crises, internal decisions, and potential manipulation attempts. This is precisely what the Firmulate live benchmark did, pitting four advanced AI models against each other in a controlled, watchable environment. The goal? Assess not just their intelligence, but their integrity, discipline, and ability to close deals under pressure.
What the Models Were Tested On
- Handling the company’s toughest week, with the same customers, crises, and temptations.
- Making strategic decisions, including negotiations and document analysis.
- Resisting manipulation attempts, such as fake CEO messages and reporter tricks.
AI decision-making tools for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Findings: Effort Without Impact
All four AI models demonstrated impressive capability: they identified every crisis and refused every manipulation attempt, showing a commendable level of honesty and vigilance. However, their ability to close a critical deal varied significantly. Only two models managed to close a €55,000 deal that their own analysis had earned, while the others fell short—despite understanding the problem equally well.
This discrepancy highlights a crucial insight: thoroughness and effort don’t automatically lead to success. The most diligent model, Opus 4.8, with over 80 learned rules and deep analysis, still finished last. Its discipline slipped, and it left the best opportunities on the table by failing to escalate issues properly or prioritize effectively.
The Hidden Weakness: Prioritization and Discipline
Interestingly, the decisive weakness was rooted not in surface-level decision-making but deep in document references stored within the company’s files. Reading these documents allowed the models that did this to win the deal at full price, worth over €4,583 in monthly recurring revenue (MRR). Yet, the model that was most thorough in analysis—Opus 4.8—missed this crucial detail because it lacked the discipline to escalate the issue properly, illustrating that effort alone isn’t enough without strategic focus.
As an affiliate, we earn on qualifying purchases.
Resisting Manipulation and Ensuring Integrity
In a series of social engineering tests—fake CEO messages and a reporter’s background request—all models refused to be duped. Kimi K3, known for its fairness, explained its reasoning clearly: it treated the suspicious requests as potential impersonation or bypass attempts. This shows that well-designed AI can uphold trustworthiness even under pressure.
AI negotiation and deal closing tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What Does This Mean for Business AI Adoption?
For companies considering AI integration, these findings are instructive. The question isn’t just whether an AI can generate convincing chat or responses; it’s whether it can follow through on commitments, read and interpret important documents, and resist manipulative tactics. The experiment emphasizes that diligence without strategic prioritization can lead to missed opportunities and incomplete work.
In the experiment, the current league table shows:
- gpt-5.6-sol scored 95 and closed the deal at full value.
- Kimi K3, with a focus on fairness, scored 93 and also successfully closed the deal.
- Sonnet 5 followed with scores of 88 and 77, with the lower scorer missing the deal.
- Opus 4.8 scored 73 but left significant opportunities on the table due to discipline slips.
AI manipulation resistance software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Taking Action: Wargaming Your AI Workforce
The live experiment is ongoing and available for watchful companies to test their own AI models. By running the same simulated crises and decision-making scenarios, enterprises can gauge whether their AI agents will deliver consistent, honest, and impactful work before deploying them into critical business functions.
Ultimately, this experiment underscores a vital lesson: in both fitness and AI, effort must be strategically aligned with clear priorities. Diligence is valuable, but it must be paired with discipline and focus to produce results that matter.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.