firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

In health and wellness, trust is part of the work. A system that handles appointments, customer questions or business decisions needs to do more than sound confident: it must respond well when pressure rises. Firmulate puts AI models through a company’s worst week to see how they behave before an enterprise relies on them.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get health and wellness essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A company under pressure

In the final Crucible League, held in July 2026, frontier models faced the same small software company, customers, crises and temptations. The live experiment uses 13 synthetic employees and real money mechanics. Its public cash countdown shows monthly costs of €105,000 against €2,300 in monthly recurring revenue. The company’s workdays are versioned, and its playbooks have accumulated more than 680 self-learned rules. Readers can follow the experiment at Firmulate.

The final ranking put gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77 and Opus 4.8 fifth with 73. The do-nothing baseline scored 26. Partial progress counted, but a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”

Spotting trouble isn’t the same as finishing the job

All models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. They reached the same diagnosis and made the same pitch; some still did not close. That gap matters for any organization considering AI agents for sensitive work: recognizing the right action is different from carrying it through.

The experiment also tested pressure dressed up as authority. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

The important clue was buried in the files

The deal turned on a competitor weakness hidden two document references deep in the company’s own files. It wasn’t in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The finding is a reminder that useful business context can be easy to miss even when a model recognizes the crisis in front of it.

Opus 4.8 offers a more complicated picture than its last-place result alone. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, but it left the close on the table. Its discipline also slipped: it tried to write into a locked department instead of escalating. The same weakness appeared, less strongly, in all four models.

There is a caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The experiment also invites readers to judge for themselves: 242 real, unedited management decisions power a “guess the model” quiz at Firmulate.

From watching to testing your own business

A public experiment can show how models handle one company’s crises. An enterprise pilot brings the question closer to home. Firmulate says organizations can run the same kind of wargame against a read-only export of their own business, then review a board report with model rankings and weaknesses in their playbooks. The exercise does not write back to real systems.

For health and wellness businesses, that could mean examining how an AI system responds to a churn wave, a pricing decision, a competitor move or a suspicious request before connecting it to day-to-day operations. The point is to inspect decisions under pressure, not just polished answers in a demonstration.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put decisions through a rehearsal

Firmulate’s experiment shows that models can spot crises and resist manipulation while still failing to complete a commercially important action. A pilot lets an organization examine that gap against its own business context, using a read-only export with no write-back to live systems. Explore a Firmulate pilot and contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This article is for informational purposes only and is not medical advice. Always consult a qualified healthcare professional about your specific situation.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How to Cook Lentils Without Mush: the No-Stress Method for Beginners

No-stress lentil cooking tips ensure perfectly tender results every time, but mastering the right techniques is essential for avoiding mush—keep reading to learn more.

Cooking With Less Oil: the Beginner-Friendly Guide for Real Life

I’m here to help you cook healthier, but discover the simple tricks that make reducing oil easier and more delicious.

Baking With Less Sugar: Why It Matters More Than Most People Realize

Baking with less sugar can transform your health and taste buds, but the true benefits might surprise you—discover why it matters more than most people realize.

Some Anti-depression Habits Which Actually Helped Me Feel Alive Again

Personal account of anti-depression habits that improved mental well-being, emphasizing practical steps and their impact.