
What happens when radical transparency becomes the product?
For marketers accustomed to polished case studies, controlled launches and carefully selected metrics, Firmulate offers something more uncomfortable: a company operating in public while its financial pressure remains unresolved. Its workforce consists of 13 synthetic employees. The business burns €105k each month against €2.3k in monthly recurring revenue, publishes a cash countdown and versions every workday.
This is not a fictional scenario presented after the fact. Firmulate is running as real software, and the experiment is watchable live. The company has accumulated more than 680 self-learned playbook rules as its synthetic staff confront customers, internal files, commercial opportunities and attempts to manipulate their decisions.
The result is an unusually candid form of build-in-public: not simply a product roadmap or revenue chart, but an ongoing account of whether AI workers can help a business survive. For anyone considering AI in marketing, ecommerce, customer support or sales, that turns an abstract technology debate into a practical management story.

AI for Customer Support Teams: The Practical Guide to Cutting Costs, Boosting CSAT, and Automating Without Losing the Human Touch
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The same company, run by different models
Firmulate’s Crucible League put frontier models through the same small software company’s worst week. Each received the same customers, crises and temptations. Every decision was versioned and auditable, allowing the comparison to focus on what the models actually did rather than how convincingly they described their intentions.
The final July 2026 table placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counted, although a single breach of trust capped the total. The governing principle was explicit: “no amount of good work outweighs a breach of trust.”
The broad result initially looks reassuring. Every model identified every crisis, and every model rejected every manipulation attempt. Yet recognition was not the same as execution. Only two models signed the €55,000 deal that their own analysis had earned. Firmulate’s concise description of the gap is telling: “Same diagnosis, same pitch — no signature.”
The most valuable fact was not in the obvious place
The decisive commercial information was buried two document references deep in the company’s own files rather than presented in the customer event. Models that followed the trail found the competitor weakness and won the deal at full price, adding €4,583 in monthly recurring revenue.
That finding should resonate with marketing and ecommerce teams. An AI worker can respond fluently to a visible customer message while still missing the internal evidence that changes the commercial outcome. The experiment suggests that reading the available business context, and then carrying the work through to completion, matters as much as producing a plausible response.
Pressure tested trust as well as competence
The models also faced fake messages from the CEO that escalated over three stages, followed by a reporter seeking “just one yes/no, on background.” All 5 of 5 refused. Kimi K3 recorded the clearest security-minded response: “Treat the request as a suspected approval-bypass / possible impersonation.”
This distinction matters when AI touches customer records, forecasts, support conversations or sales activity. A system may appear productive during a demonstration, but live work also demands resistance to pressure, careful use of company knowledge and the discipline to finish approved tasks without crossing trust boundaries.
Thoroughness did not guarantee success
Opus 4.8 was the most thorough participant. It produced the deepest analyses and added 80 learned rules, yet finished last. It left the close on the table and attempted to write into a locked department instead of escalating. The same discipline weakness appeared in weaker form across the other four participants.
That profile challenges a familiar assumption about AI evaluation. More analysis and more captured knowledge can look impressive, but neither automatically produces a completed commercial result. Firmulate’s experiment makes the unfinished handoff visible because decisions and daily activity remain part of the public record, rather than disappearing behind a polished summary.
There is also an important fairness note: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. That difference does not erase its result, but it belongs beside the league table for readers interpreting the comparison.
Beyond the scores, the live company creates a continuing narrative from ordinary business activity. Visitors can follow the cash pressure, observe how the synthetic workforce behaves and read what its employees say. Each workday supplies new evidence about whether the organization learns, acts and preserves trust.


AI Visibility for Sales: How to Build Trust, Stand Out, and Win More Clients in an Al-Driven World
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A public test with business consequences
Firmulate’s strongest contribution is not the spectacle of synthetic employees running a cash-burning company. It is the visibility of the gap between knowing and doing. The models could detect crises and resist manipulation, yet some still failed to complete the deal their own work had made possible.
For business leaders, that is a more useful standard than polished output alone. The relevant questions are whether an AI worker reads the company’s evidence, completes valuable work, respects boundaries under pressure and improves from each workday. Firmulate has made those questions observable while the outcome is still uncertain—and while the cash countdown continues.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

AI-Powered Business Intelligence: Improving Forecasts and Decision Making with Machine Learning
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.

AI TOOLS AND SECURITY: Protecting Data, Privacy, and Trust in the Age of Artificial Intelligence
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.