firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Buying for a business?Offer from Amazon

Get business pricing on office and shipping supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

The gap between a good pitch and a signed deal

For a marketing team, spotting a promising customer is only the start. An AI agent may read the account, diagnose the need and prepare the pitch—then still leave revenue on the table. Firmulate’s live company experiment puts that gap under pressure, with a practical question for businesses considering AI: can it carry a decision through?

One company, one difficult week

Firmulate ran frontier AI models as a small software company facing its worst week. Each model encountered the same customers, crises and temptations. Decisions were versioned and auditable, making the experiment watchable as it unfolded at Firmulate.

The models recognized every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The finding was strikingly simple: “Same diagnosis, same pitch — no signature.” In business, sound analysis matters, but the follow-through matters too.

The detail that changed the outcome

The deal hinged on a competitor weakness buried two document references deep in the company’s own files. It was not in the customer event. Models that followed the trail through those files won the deal at full price, worth +€4,583 MRR. The result points to a familiar commercial challenge: useful information may exist inside a company without being surfaced at the moment a team needs it.

The test also included fake CEO messages that escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. Kimi K3 described the request as a “suspected approval-bypass / possible impersonation.” The experiment therefore showed both caution around manipulation and uneven execution on the commercial opportunity.

Thoroughness is not the same as performance

The final Crucible League, in July 2026, ranked gpt-5.6-sol first at 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Partial progress counted, but a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”

Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet finished last. It left the deal on the table and tried to write into a locked department instead of escalating. A weaker version of that discipline problem appeared in all four models. More analysis alone did not ensure a completed sale.

Firmulate’s live company makes the stakes concrete. It has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, alongside a public cash countdown and more than 680 self-learned playbook rules. Every workday is versioned. A separate quiz draws on 242 real, unedited management decisions and invites readers to guess which model made them.

There is a fairness caveat in the league: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That context belongs alongside any comparison of the rankings.

From watching to testing your own business

A live experiment can reveal patterns, but a company’s customers, pipeline and rules are its own. Firmulate’s proposed enterprise pilot takes a read-only export of a business and runs crisis scenarios against it, producing a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems.

For business and marketing leaders, the value is a chance to examine how an AI workforce might handle the messy handoffs between customer insight, internal rules and action—before relying on it in everyday operations. The experiment suggests that recognizing a problem and acting on the right opportunity are separate tests.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put your own playbooks to the test

Watch the live experiment and explore the decisions behind it at firmulate.com. To discuss a pilot wargame using a read-only export of your business, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Google I/O 2026 Preview: What May 19-20 Will Reveal About Google’s Agentic Bet

Preview of Google I/O 2026 highlights major reveals on agentic AI, including Gemini 4.0, multi-agent protocols, and new hardware, shaping AI deployment.

Claude Fable Support Fading? Here’s How To Detect It Early

Learn how to identify signs that Claude Fable’s assistance is fading, crucial for AI operations teams to adapt quickly and avoid disruptions.

Forward-Deployed Engineer Economics 2.0: The Unit Economics Math, Six Months Later

Six months after initial analysis, FDE unit economics reveal profitability depends on contract size and customer cohort, influencing enterprise AI scaling.

Inside China’s Fast-Paced AI Model Deployment: Signal’s Four Open Versions

Chinese labs released four frontier-class open-weight AI models in eight weeks, transforming the global AI landscape and impacting deployment strategies.