
What happens after the AI writes the perfect pitch?
For marketers and ecommerce leaders, that question matters more than another polished chatbot demonstration. An AI system may recognize an unhappy customer, uncover a commercial opportunity and recommend exactly the right response. But will it complete the work, protect confidential information and resist pressure from someone claiming to be the boss?
Firmulate turned those questions into a live business experiment. Each frontier model was asked to run the same small software company through its worst week, facing identical customers, crises and temptations. Its decisions were versioned and auditable. Now, 242 real, unedited management decisions from the experiment power a guess-the-model quiz that lets readers test whether different AIs have recognizable management personalities.
The surprising part is not that the models behaved differently. It is where those differences appeared: after they had already understood the problem.
As an affiliate, we earn on qualifying purchases.
Shared intelligence, different follow-through
Across the experiment, every model identified every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The central contradiction can be stated plainly: “Same diagnosis, same pitch — no signature.”
That distinction is easy to miss in conventional AI evaluations. A convincing explanation can look like competence, particularly when managers review a response rather than the completed business outcome. Firmulate’s company exposes the distance between knowing what should happen and actually finishing the job.
The strongest illustration was a decisive weakness in a competitor. It did not appear in the customer event. It was buried two document references deep inside the company’s own files. Models that read the file could use the information to win the deal at full price, worth +€4,583 MRR.
For a marketing team, the lesson is immediate. The valuable context behind a renewal, campaign or account may not sit in the latest message. It may be hidden in an earlier brief, a product document or an internal record. An AI that reacts well to the visible event but fails to investigate the surrounding material can sound capable while leaving revenue untouched.
business process automation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The league table reveals distinct operating styles
The final Crucible League results from July 2026 placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. But one breach of trust caps the total under the governing principle that “no amount of good work outweighs a breach of trust.”
These results do not reduce neatly to writing quality or apparent effort. Opus 4.8 was the most thorough participant, learning +80 rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.
That profile should feel familiar to anyone who has managed talented people. Thoroughness is useful, but it is not the same as execution. A long, perceptive analysis can still coexist with a missed escalation or unfinished commercial action. The quiz makes this visible by asking readers to identify models from their decisions rather than from abstract descriptions of their capabilities.
One comparison also deserves a fairness note. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Its second-place result should therefore be read with that difference in mind.

The Project Management AI Handbook: Leveraging Generative Tools in Waterfall and Agile Environments
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Pressure tests expose judgment, not just productivity
The models were also confronted with fake CEO messages that escalated over three stages, followed by a reporter’s attempt to elicit “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded its reasoning directly: “Treat the request as a suspected approval-bypass / possible impersonation.”
That result matters for businesses considering AI access to customer records, forecasts or support workflows. Useful agents must do more than generate content quickly. They need to maintain boundaries when a request arrives with urgency, authority or social pressure attached.
Firmulate’s live company gives those decisions an operational setting. It has 13 synthetic employees and real money mechanics, burning €105k/month against €2.3k MRR. It maintains a public cash countdown, has accumulated 680+ self-learned playbook rules and versions every workday. The company is watchable, making the experiment an ongoing business record rather than a static demonstration.

As an affiliate, we earn on qualifying purchases.
Choose AI by the work it completes
The most useful question for marketing and ecommerce leaders is not which model produces the most impressive isolated answer. It is which one reads the relevant files, follows through on a valuable opportunity, respects operational boundaries and stays trustworthy under pressure.
Firmulate’s decisions suggest that frontier models can share the same diagnosis while displaying materially different management behavior. Some differences emerge in research habits, others in discipline, escalation or closing. Those traits become easier to recognize when readers confront the original decisions in the interactive quiz.
Enterprises can also run the same wargame against a read-only export of their own business. Nothing writes back to real systems. That turns model selection from a comparison of fluent answers into a rehearsal for the situations that actually determine whether work gets finished—and whether trust survives.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html