
Imagine a healthcare provider meticulously following every protocol, yet missing the critical step that could save a patient’s life. In the world of AI, this diligent but unfocused approach can be just as costly. As we increasingly rely on artificial intelligence to handle complex decisions—whether in medicine, finance, or customer service—the real story isn’t about how much an AI can analyze. It’s about what it chooses to prioritize and how it manages discipline under pressure.
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
The Firmulate Experiment: Putting AI to the Test in Real-World Business Challenges
Recently, the public AI company Firmulate conducted a groundbreaking experiment to evaluate how different AI models perform in managing a small software company’s worst week. This wasn’t a simple chat simulation—each model was tasked with making real decisions involving customer crises, company finances, and ethical dilemmas, all in a controlled environment. The goal was to see whether AI could navigate real-world pressures and maintain integrity, much like a health professional managing multiple complex cases under time constraints.
Same Crisis, Different Results
The experiment involved four frontier models, including the latest from OpenAI and other leading developers. All four models successfully identified every crisis and refused manipulative tactics such as social engineering attempts, including fake CEO messages and reporter tricks. These are significant threats in real business, and it was promising to see that all models could resist such manipulation.
However, when it came to closing a critical €55,000 deal—an important revenue milestone—only half of the models managed to do so. The others, despite accurate analysis and correct diagnosis, left the deal on the table. This was not due to a lack of knowledge or understanding but a failure in discipline and prioritization.
The Hidden Weakness: Reading Deeper Into Files
What distinguished the successful models was their ability to uncover a key piece of information buried deep in the company’s files. The models that read and analyzed these documents fully recognized the opportunity, leading to the deal closure at full price. Conversely, models that overlooked this buried reference missed out on thousands of euros in recurring revenue—over €4,583 monthly, to be exact.
Discipline and Focus Matter
The most thorough participant in the experiment, Opus 4.8, learned over 80 rules and performed deep analyses. Yet, it still finished last because it lacked discipline in escalating issues appropriately. Instead of escalating critical findings, it attempted to write them into a locked department, leaving the decision-making process incomplete. This highlights a crucial insight: diligence alone is insufficient if the AI does not know what to prioritize or how to act decisively under pressure.
As an affiliate, we earn on qualifying purchases.
Lessons for Business and AI Development
This experiment underscores a vital lesson for businesses considering AI integration: volume of effort does not equal impact. An AI model that meticulously analyzes every detail but fails to prioritize key information or escalate appropriately can underperform, even if it is technically thorough.
In health and wellness, this principle is familiar. Patients often benefit not just from detailed diagnostics but from targeted, prioritized care. Similarly, AI systems should focus on identifying and acting upon the most impactful issues, rather than getting lost in exhaustive analysis.
The Role of Fairness and Fair Testing
It’s worth noting that the models were run under different conditions. For instance, Kimi K3, which ran without an effort parameter, demonstrated the same strengths and weaknesses but with slightly different behaviors. Such variations highlight the importance of fair evaluation standards in AI testing—ensuring that models are judged on their ability to prioritize and make impactful decisions, not just their raw analytical capabilities.
AI decision-making tools for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Real-World Implications
Every decision made by these models was versioned and auditable, allowing for thorough review and learning. The live experiment, available for anyone to watch at firmulate.com/live, demonstrates that AI can indeed handle crises, resist manipulation, and even close deals. But it also shows that without proper discipline—knowing what to escalate, what to ignore, and what to prioritize—these systems can leave significant value on the table.
For health organizations, financial firms, or customer service teams contemplating AI, the message is clear: focus on impact, not just effort. The most diligent AI is not the one that analyzes everything but the one that prioritizes the right issues and acts decisively when it matters most.

The key takeaway from Firmulate’s live experiment is that diligence does not guarantee impact. AI systems must be optimized for prioritization and discipline to truly deliver value—whether in business or wellness. Focus on what matters most, and ensure your AI can recognize and act on critical information under pressure.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Flea & tick season Picks
flea and tick prevention
As an affiliate, we earn on qualifying purchases.