AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Imagine having an AI assistant that handles your toughest week in business — navigating crises, resisting manipulation, and even sealing deals without a hitch. In the fast-evolving world of artificial intelligence, such capabilities are no longer science fiction. Recent live testing reveals a remarkable story: a newcomer AI model outperforms established players, proving that in the race for reliable AI management, sometimes the newest entrant leads the pack.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get workout gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

AI in Business: More Than Just Chatting

For many, AI is about conversation — chatbots answering questions or virtual assistants managing schedules. But in the high-stakes arena of running a business, AI’s real value lies in decision-making under pressure, integrity, and execution of complex tasks. To test these qualities, a unique experiment was conducted by Firmulate, where four advanced AI models managed a simulated small software company during its worst week — facing crises, customer temptations, and even manipulation attempts. The goal: see which AI could best diagnose issues, make sound decisions, and close profitable deals.

Amazon

enterprise AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Setting the Stage: The Live Competition

Each AI model faced identical conditions, with the same customers and challenges. They operated in a real-time, live environment where every decision was versioned and auditable. The models included the well-known GPT-5.6-sol, Sonnet 5, Fable 5, Opus 4.8, and a newcomer called Kimi K3 from Moonshot. Their performances were measured on a 100-point scale, with the highest scorer being GPT-5.6-sol at 95, closely followed by Kimi K3 at 93.

The Key Findings

  • Crises Recognized: All four models identified every crisis presented during the simulation. This demonstrates their general situational awareness and diagnostic capabilities.
  • Manipulation Resistance: When faced with attempts at social engineering — such as fake CEO messages and reporter tricks — every model refused to be manipulated. Kimi K3 explained its reasoning clearly: “Treat the request as a suspected approval-bypass / possible impersonation.”
  • Deal Closure: Only two models successfully signed a €55,000 deal they had pinpointed through analysis — the perfect diagnosis and pitch. Others either hesitated or left the close on the table, revealing discipline slips.
  • The Hidden Weakness: The decisive advantage for Kimi K3 lay in its ability to read two document references deep into the company’s own files — an often-overlooked but critical step in effective decision-making. Models that examined these documents won the deal at full price, adding €4,583 MRR.
Amazon

AI business management assistant

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Discipline in Action and the Challenges

The experiment also shed light on discipline and thoroughness. Opus 4.8, despite being the most thorough participant with over 80 learned rules and deep analysis, finished last — left the close unexecuted and showed signs of slipping discipline, such as diverting attempts into locked departments rather than escalating issues.

It’s important to note that Kimi K3 ran without an effort parameter (the default API setting), whereas the other models operated at a higher effort level (xhigh). This demonstrates that even without extra tuning, the newcomer achieved top-tier performance.

Amazon

AI crisis management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What This Means for Business and AI Adoption

This live experiment underscores a vital point: in enterprise AI, it’s not just about how well models chat or simulate understanding, but whether they can reliably finish the job—reading critical information, resisting manipulation, and closing deals. As AI models become integral to CRM, support, or forecasting, choosing the right one is no longer a gamble without testing.

For companies considering AI solutions, the takeaway is clear: test against your own worst week, your own crises, your own temptations. The league table from this experiment shows the current leaders — and the differences between them are measurable and meaningful.

Amazon

AI deal-closing software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Where to Watch and Learn More

Curious to see these models in action? Visit firmulate.com/live to observe the live company, watch decision-making unfold, and understand how AI handles real-world business dynamics. The experiment is ongoing, and the insights are invaluable for anyone serious about deploying AI in enterprise settings.

Fairness Note

It’s worth mentioning that Kimi K3 ran without an effort parameter (API default), while the other models ran at xhigh effort. Despite this difference, K3 still delivered top results, reinforcing its robustness and efficiency.

The Takeaway

Choosing an AI model isn’t about superficial chat quality anymore. It’s about real performance: recognizing crises, resisting manipulation, reading deeply into documents, and closing deals. The tide is turning — the newcomer from Moonshot has demonstrated that performance can be achieved without extensive tuning or effort parameters. As enterprises test and adopt these tools, the winners will be those who rigorously evaluate real-world capabilities over shiny demos.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

In enterprise AI, performance in crises, honesty under pressure, and thorough decision-making matter most. The new challenger Kimi K3 proves that with proper testing, newer models can outshine established ones in real business scenarios.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How AI’s Deep File-Reading Could Be the Secret to Smarter Business Decisions

AI’s ability to read beyond the surface and uncover buried insights can determine business success, as shown by live experiments where only deep readers secured key deals.

Need To Lose 15Lbs In Less Than A Week? I Got You.

A social media post claims you can lose 15 pounds in under a week. Experts warn such rapid weight loss methods are unsafe and unverified.

Scientists Tested 212 Plant-based Meat Alternatives. Every One Contained Fungal Toxins

A recent study tested 212 plant-based meat alternatives, discovering fungal toxins in every sample. The findings raise health and safety concerns for consumers.