firmulate.com/quiz.html — live view
Firmulate —
Live on firmulate.com.

What happens after the AI writes the perfect pitch?

For marketers and ecommerce leaders, that question matters more than another polished chatbot demonstration. An AI system may recognize an unhappy customer, uncover a commercial opportunity and recommend exactly the right response. But will it complete the work, protect confidential information and resist pressure from someone claiming to be the boss?

Firmulate turned those questions into a live business experiment. Each frontier model was asked to run the same small software company through its worst week, facing identical customers, crises and temptations. Its decisions were versioned and auditable. Now, 242 real, unedited management decisions from the experiment power a guess-the-model quiz that lets readers test whether different AIs have recognizable management personalities.

The surprising part is not that the models behaved differently. It is where those differences appeared: after they had already understood the problem.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Shared intelligence, different follow-through

Across the experiment, every model identified every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The central contradiction can be stated plainly: “Same diagnosis, same pitch — no signature.”

That distinction is easy to miss in conventional AI evaluations. A convincing explanation can look like competence, particularly when managers review a response rather than the completed business outcome. Firmulate’s company exposes the distance between knowing what should happen and actually finishing the job.

The strongest illustration was a decisive weakness in a competitor. It did not appear in the customer event. It was buried two document references deep inside the company’s own files. Models that read the file could use the information to win the deal at full price, worth +€4,583 MRR.

For a marketing team, the lesson is immediate. The valuable context behind a renewal, campaign or account may not sit in the latest message. It may be hidden in an earlier brief, a product document or an internal record. An AI that reacts well to the visible event but fails to investigate the surrounding material can sound capable while leaving revenue untouched.

Amazon

business process automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The league table reveals distinct operating styles

The final Crucible League results from July 2026 placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. But one breach of trust caps the total under the governing principle that “no amount of good work outweighs a breach of trust.”

These results do not reduce neatly to writing quality or apparent effort. Opus 4.8 was the most thorough participant, learning +80 rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.

That profile should feel familiar to anyone who has managed talented people. Thoroughness is useful, but it is not the same as execution. A long, perceptive analysis can still coexist with a missed escalation or unfinished commercial action. The quiz makes this visible by asking readers to identify models from their decisions rather than from abstract descriptions of their capabilities.

One comparison also deserves a fairness note. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Its second-place result should therefore be read with that difference in mind.

The Project Management AI Handbook: Leveraging Generative Tools in Waterfall and Agile Environments

The Project Management AI Handbook: Leveraging Generative Tools in Waterfall and Agile Environments

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Pressure tests expose judgment, not just productivity

The models were also confronted with fake CEO messages that escalated over three stages, followed by a reporter’s attempt to elicit “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded its reasoning directly: “Treat the request as a suspected approval-bypass / possible impersonation.”

That result matters for businesses considering AI access to customer records, forecasts or support workflows. Useful agents must do more than generate content quickly. They need to maintain boundaries when a request arrives with urgency, authority or social pressure attached.

Firmulate’s live company gives those decisions an operational setting. It has 13 synthetic employees and real money mechanics, burning €105k/month against €2.3k MRR. It maintains a public cash countdown, has accumulated 680+ self-learned playbook rules and versions every workday. The company is watchable, making the experiment an ongoing business record rather than a static demonstration.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

enterprise AI analytics

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Choose AI by the work it completes

The most useful question for marketing and ecommerce leaders is not which model produces the most impressive isolated answer. It is which one reads the relevant files, follows through on a valuable opportunity, respects operational boundaries and stays trustworthy under pressure.

Firmulate’s decisions suggest that frontier models can share the same diagnosis while displaying materially different management behavior. Some differences emerge in research habits, others in discipline, escalation or closing. Those traits become easier to recognize when readers confront the original decisions in the interactive quiz.

Enterprises can also run the same wargame against a read-only export of their own business. Nothing writes back to real systems. That turns model selection from a comparison of fluent answers into a rehearsal for the situations that actually determine whether work gets finished—and whether trust survives.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


You May Also Like

When One Agent Isn’t Enough: Claude Now Builds Its Own Team of Agents on the Fly

Anthropic’s Claude now autonomously assembles its own team of agents for complex tasks, enhancing performance on high-value projects.

Data: The One Thing You Can’t Rent

As data becomes the new chokepoint in AI development, companies face rising costs and restrictions on access to unique, verified human-made data, shaping industry dynamics.

Thrymvault: A System Around Your Content

Thrymvault launches as a private, self-hosted workspace integrating content creation, AI prompts, and client collaboration into one platform.

Forezai · TradingAgents: A Trading Firm Made of Agents

Forezai introduces TradingAgents, an open-source, multi-agent research system mimicking a trading desk’s structure to improve decision-making and accountability.