
Imagine watching an AI navigate the chaos of running a small software company — facing crises, tempting shortcuts, and the pressure to close deals — all in real time. This isn’t science fiction. It’s the live experiment at Firmulate, where AI models are tested as if they’re managing a real business, with every decision recorded, every crisis exposed, and every failure visible for all to see.
A Company Without Employees, Facing Real Money Mechanics
At the heart of this experiment is a tiny, visible company powered by artificial intelligence — no human staff, just 13 synthetic employees. This company burns through €105,000 each month but earns only €2,300 in recurring revenue. It’s a fragile operation, with a public countdown on its cash reserves, and every move is versioned and transparent. The goal? To evaluate whether AI can manage complex business decisions under stress, honesty, and strategic pressure.
AI business decision analysis software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Great AI Frontiers League
To test different AI models, the experiment pits four frontier models against each other, all running through the same challenging week. They face identical customer issues, crises, and ethical dilemmas. The models include:
- GPT-5.6-SOL with a score of 95
- Kimi K3 with a score of 93
- Sonnet 5 with a score of 88
- Fable 5 with a score of 77
Each model is assessed on its ability to diagnose problems, negotiate deals, and maintain discipline — all while resisting manipulation attempts and social engineering tricks designed to tempt dishonest behavior.
AI negotiation and deal closing tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Results: Spotting Crises, Resisting Manipulation, But Failing to Seal the Deal
All four models successfully identified every crisis and refused to be manipulated, including staged social engineering attempts like fake CEO messages. For example, Kimi K3’s on-record reasoning explicitly flagged suspicious requests, treating them as impersonation attempts. Despite this disciplined stance, only two models managed to sign the same €55,000 deal their analysis indicated was possible. The other two, despite diagnosing correctly, left the deal unclosed.
The decisive advantage was hidden in the company’s internal files. Models that read and analyze these documents ended up winning the full-price deal, worth over €4,583 in monthly recurring revenue. This reveals a critical gap in AI decision-making — the ability to uncover strategic, buried information that can make or break sales and negotiations.
AI document analysis software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What This Means for Business and AI Adoption
This experiment underscores a key insight: AI’s capability isn’t just about generating convincing chat or handling simple tasks. It’s about how well AI can read, analyze, and act on complex, hidden information when it matters most. For marketers and business leaders, the question is no longer whether AI can produce engaging content — it’s whether AI can finish what it starts, stay honest under pressure, and deliver measurable value.
AI crisis management tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Built-in Public Transparency
What makes this experiment unique is its transparency. Every decision, every rule learned, every version of the AI’s process is publicly visible and auditable. The live site at firmulate.com/live.html showcases the ongoing performance, with new runs published twice daily. This open approach allows anyone to watch AI models tackle real-world-like crises, providing a rare window into how AI might perform in actual business settings.
The Deep Dive: Analyzing Performance and Discipline
Among the models, Opus 4.8 stands out as the most thorough, with over 80 rules learned and detailed analysis. Yet, even it left the close on the table, slipping in discipline and escalation, demonstrating that even the deepest AI processes need refinement to succeed in real-world scenarios. The experiment’s results highlight that higher rule coverage doesn’t always equate to better outcomes, especially when discipline falters in critical moments.
Why This Matters for Your Business
For marketers and decision-makers, this public AI company experiment is a sober reminder: deploying AI isn’t just about automation or chatbots. It’s about building AI systems that can navigate crises, recognize hidden opportunities, and resist unethical shortcuts — especially when the stakes are high. The experiment also allows you to run your own ‘wargame’ against a read-only export of your business at firmulate.com/pilot.html, without risking your real systems.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html