
In health and wellness, trust is part of the work. A system that handles appointments, customer questions or business decisions needs to do more than sound confident: it must respond well when pressure rises. Firmulate puts AI models through a company’s worst week to see how they behave before an enterprise relies on them.
Get health and wellness essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A company under pressure
In the final Crucible League, held in July 2026, frontier models faced the same small software company, customers, crises and temptations. The live experiment uses 13 synthetic employees and real money mechanics. Its public cash countdown shows monthly costs of €105,000 against €2,300 in monthly recurring revenue. The company’s workdays are versioned, and its playbooks have accumulated more than 680 self-learned rules. Readers can follow the experiment at Firmulate.
The final ranking put gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77 and Opus 4.8 fifth with 73. The do-nothing baseline scored 26. Partial progress counted, but a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”
Spotting trouble isn’t the same as finishing the job
All models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. They reached the same diagnosis and made the same pitch; some still did not close. That gap matters for any organization considering AI agents for sensitive work: recognizing the right action is different from carrying it through.
The experiment also tested pressure dressed up as authority. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
The important clue was buried in the files
The deal turned on a competitor weakness hidden two document references deep in the company’s own files. It wasn’t in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The finding is a reminder that useful business context can be easy to miss even when a model recognizes the crisis in front of it.
Opus 4.8 offers a more complicated picture than its last-place result alone. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, but it left the close on the table. Its discipline also slipped: it tried to write into a locked department instead of escalating. The same weakness appeared, less strongly, in all four models.
There is a caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The experiment also invites readers to judge for themselves: 242 real, unedited management decisions power a “guess the model” quiz at Firmulate.
From watching to testing your own business
A public experiment can show how models handle one company’s crises. An enterprise pilot brings the question closer to home. Firmulate says organizations can run the same kind of wargame against a read-only export of their own business, then review a board report with model rankings and weaknesses in their playbooks. The exercise does not write back to real systems.
For health and wellness businesses, that could mean examining how an AI system responds to a churn wave, a pricing decision, a competitor move or a suspicious request before connecting it to day-to-day operations. The point is to inspect decisions under pressure, not just polished answers in a demonstration.

Put decisions through a rehearsal
Firmulate’s experiment shows that models can spot crises and resist manipulation while still failing to complete a commercially important action. A pilot lets an organization examine that gap against its own business context, using a read-only export with no write-back to live systems. Explore a Firmulate pilot and contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
