
Imagine a Company Running in Real Time—Without Humans or Profit—Yet Fully Transparent
For those interested in health and wellness, transparency and resilience are key. Now, imagine a company that’s entirely built and operated by AI, with no employees, yet constantly battling its own financial and operational crises in full public view. Welcome to the world of Firmulate, where the extraordinary becomes reality—a live experiment revealing just how far AI can go in managing real-world business challenges.

Building AI-Powered Products: The Essential Guide to AI and GenAI Product Management
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Firmulate Experiment: A Glimpse Into AI-Driven Business Management
At the heart of this experiment is a small software company run entirely by AI models, watched closely by the public. Every weekday, the system faces dozens of crises—customer issues, internal miscommunications, ethical dilemmas—and makes decisions, all while being watched by thousands online. The company employs 13 synthetic employees, each governed by over 680 self-learned rules, and operates in real time at a monthly burn rate of €105,000 against a revenue of just €2,300.
This setup isn’t just a tech demo; it’s a rigorous test of AI’s ability to manage complex, unpredictable, and ethically fraught situations. Every decision made by these models is logged, versioned, and publicly accessible at firmulate.com/live.html, making it a unique transparency experiment.
Performance in Crisis and Deal-Making
Four frontier AI models have been pitted against each other, each running the same tough week. Their goal: navigate customer crises, avoid manipulation, and secure business deals. Remarkably, all models identified every crisis and refused every unethical or manipulative attempt, even when faced with social engineering tactics like fake CEO messages or reporter tricks. Only two models managed to close a €55,000 deal, which was earned through accurate diagnosis and appropriate pitches. The others failed to sign, despite knowing the opportunity existed.
The key insight lay in the models’ ability to read and understand internal company documents—something that isn’t visible in typical chat demos. The models that read deeper into the files uncovered critical weaknesses that led to successful deal closures, translating into over €4,583 in monthly recurring revenue.
Operational Challenges and Weaknesses
Despite their intelligence, the models showed vulnerabilities. For instance, the most thorough participant, OPUS 4.8, with over 80 learned rules, left deals on the table and showed lapses in discipline—writing attempts into a locked department rather than escalating them. This pattern was consistent across all models, revealing a fundamental challenge: AI decision-making under stress isn’t flawless, especially when discipline and process adherence matter.
Ethics, Trust, and Real-World Risks
A notable aspect of the experiment is the AI’s unwavering refusal of social engineering attempts. All models declined fake CEO requests and background approval tricks, with Kimi K3 explicitly treating such requests as potential impersonation or approval bypass risks. This demonstrates an encouraging level of ethical resilience—an essential feature for AI systems that might someday operate in sensitive or high-stakes environments.

The AI-Driven Leader: Harnessing AI to Make Faster, Smarter Decisions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Bigger Picture: AI as a Business Manager
This ongoing live experiment is more than a curiosity. It raises critical questions for anyone interested in AI’s role in health, wellness, or business operations. If AI agents are to be integrated into tools like CRM, support systems, or forecasting, the key questions aren’t about how well they generate text or mimic human conversation—they’re about whether they can finish what they start, stay honest under pressure, and make decisions aligned with human values.
The current leaderboard places GPT-5.6 at the top, with a score of 95 out of 100, having identified the buried fact crucial to closing the deal. Kimi K3 follows closely with a score of 93, thanks to its disciplined approach. Yet even the highest performers show room for improvement, especially in maintaining discipline during stressful situations.

AI for Solo Lawyers: A Practical Guide to AI Tools that Save You Time and Grow Your Practice (AI for Professionals)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Takeaway: Transparency and Vigilance in AI-Driven Business
The Firmulate experiment is a rare peek into how AI might manage real-world companies—flaws, successes, and all. For businesses and consumers alike, the takeaway is clear: AI’s value isn’t just in generating convincing chat or predictions but in its ability to stay true, finish what it starts, and operate ethically under pressure. As this AI-run company continues its daily struggle for survival, it offers lessons on building resilient, transparent, and trustworthy AI systems.


AI: THE PERPETUAL INTERN – Its Brilliance and Failures Share the Same Root
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Takeaway
This public AI company demonstrates that genuine management involves trust, discipline, and transparency—qualities that AI models are still developing. Watching this live experiment underscores the importance of rigorous testing and ethical safeguards before deploying AI at scale in real-world operations.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html