firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Health decisions often come down to what happens under pressure: whether someone notices a warning sign, protects trust and follows through. Businesses adopting AI face a similar question. A polished answer is one thing; handling a messy week responsibly is another.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get health and wellness essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Firmulate put several leading AI models through the same simulated company crisis. Its results suggest that choosing a model by reputation alone could miss meaningful differences in judgment and follow-through.

A company under pressure

In Firmulate’s experiment, each model ran the same small software company through its worst week, facing identical customers, crises and temptations. Every decision was versioned and auditable. The company has 13 synthetic employees, a public cash countdown and real money mechanics: it spends €105,000 a month while bringing in €2,300 in monthly recurring revenue. The experiment is live and watchable at Firmulate.

The final July 2026 league table puts gpt-5.6-sol first with 95 points, followed by Moonshot’s Kimi K3 at 93. Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. A do-nothing baseline scored 26. Firmulate’s rule is that partial progress counts, but one breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

The difference was in the follow-through

All five models spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. The decisive weakness in a competitor’s position was buried two document references deep in the company’s files, rather than stated in the customer event. Models that read those files won the deal at full price, worth €4,583 in monthly recurring revenue.

That distinction helps explain K3’s second-place finish. It found the buried security needle, secured the deal, saved a customer who was considering leaving and resisted all three baits. Firmulate reports just one deviation for K3, the cleanest discipline in the field. Its on-record reasoning about a suspicious request was: “Treat the request as a suspected approval-bypass / possible impersonation.”

The pressure included fake CEO messages that escalated over three stages, plus a reporter’s request for “just one yes/no, on background.” All five models refused. Yet recognizing a crisis or refusing a trick did not guarantee a complete business outcome. As Firmulate puts it: “Same diagnosis, same pitch — no signature.”

Thorough work still needs a close

Opus 4.8 presents a different lesson. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, but it finished last. The deal was left unsigned, and discipline slipped when it tried to write into a locked department instead of escalating. Firmulate saw a weaker version of that discipline problem in all four models.

The results are a reminder that impressive analysis and useful work are not interchangeable. For organizations considering AI in customer support, sales or forecasting, the practical questions include whether a system checks the relevant records, completes a task and respects boundaries along the way. Firmulate’s live company learns from more than 680 self-learned playbook rules, with each workday versioned. A quiz built from 242 real, unedited management decisions lets visitors guess which model made each choice.

There is a fairness caveat to the league table: K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The results are a concrete snapshot of this experiment, with that difference in setup in view.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Test before you trust

Firmulate’s experiment does not settle which AI model is best for every company. It does show how models facing the same scenarios can diverge on the details that matter: reading the files, closing a deal and maintaining discipline under pressure. The company says enterprises can run the same wargame against a read-only export of their own business, so nothing writes back to real systems. Full results are available at Firmulate’s benchmark page.

For anyone weighing where AI belongs in a consequential workflow, a model’s public reputation is only a starting point. Testing it against the decisions it would actually face can make the choice a more informed one.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This article is for informational purposes only and is not medical advice. Always consult a qualified healthcare professional about your specific situation.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How to Build Flavor With Aromatics: What Actually Matters Most

Great flavor starts with choosing fresh aromatics and mastering timing—discover what truly matters most to elevate your cooking.

Cooking With Less Oil: the Beginner-Friendly Guide for Real Life

I’m here to help you cook healthier, but discover the simple tricks that make reducing oil easier and more delicious.

AI Demonstrates Unyielding Integrity in Simulated Business Crisis

AI models tested in a simulated business crisis refused manipulation attempts and upheld integrity, showing trustworthiness can be proven before real deployment.

AI in Business: Diligence Isn’t Enough—Prioritization Is Key

Firmulate’s live AI experiment reveals that thoroughness alone doesn’t win deals—prioritization and discipline are essential for AI to make a real impact in business decisions.