firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Buying for a business?Offer from Amazon

Get business pricing on office and shipping supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

The gap between a good pitch and a signed deal

For a marketing team, spotting a promising customer is only the start. An AI agent may read the account, diagnose the need and prepare the pitch—then still leave revenue on the table. Firmulate’s live company experiment puts that gap under pressure, with a practical question for businesses considering AI: can it carry a decision through?

One company, one difficult week

Firmulate ran frontier AI models as a small software company facing its worst week. Each model encountered the same customers, crises and temptations. Decisions were versioned and auditable, making the experiment watchable as it unfolded at Firmulate.

The models recognized every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The finding was strikingly simple: “Same diagnosis, same pitch — no signature.” In business, sound analysis matters, but the follow-through matters too.

The detail that changed the outcome

The deal hinged on a competitor weakness buried two document references deep in the company’s own files. It was not in the customer event. Models that followed the trail through those files won the deal at full price, worth +€4,583 MRR. The result points to a familiar commercial challenge: useful information may exist inside a company without being surfaced at the moment a team needs it.

The test also included fake CEO messages that escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. Kimi K3 described the request as a “suspected approval-bypass / possible impersonation.” The experiment therefore showed both caution around manipulation and uneven execution on the commercial opportunity.

Thoroughness is not the same as performance

The final Crucible League, in July 2026, ranked gpt-5.6-sol first at 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Partial progress counted, but a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”

Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet finished last. It left the deal on the table and tried to write into a locked department instead of escalating. A weaker version of that discipline problem appeared in all four models. More analysis alone did not ensure a completed sale.

Firmulate’s live company makes the stakes concrete. It has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, alongside a public cash countdown and more than 680 self-learned playbook rules. Every workday is versioned. A separate quiz draws on 242 real, unedited management decisions and invites readers to guess which model made them.

There is a fairness caveat in the league: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That context belongs alongside any comparison of the rankings.

From watching to testing your own business

A live experiment can reveal patterns, but a company’s customers, pipeline and rules are its own. Firmulate’s proposed enterprise pilot takes a read-only export of a business and runs crisis scenarios against it, producing a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems.

For business and marketing leaders, the value is a chance to examine how an AI workforce might handle the messy handoffs between customer insight, internal rules and action—before relying on it in everyday operations. The experiment suggests that recognizing a problem and acting on the right opportunity are separate tests.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put your own playbooks to the test

Watch the live experiment and explore the decisions behind it at firmulate.com. To discuss a pilot wargame using a read-only export of your business, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Stenvrik: News as Geography

Stenvrik launches a live news platform pinning stories to 49 global hubs on a 3D globe, offering a new geographic approach to news consumption and trend detection.

China Sphere Capability Gap, Q2 2026 Update: Five Labs, Five Strategies, One Narrowing Frontier

Five Chinese labs launched frontier-tier models within four weeks, narrowing the US-China AI capability gap in key areas, but the US still leads on top-tier tasks.

Best Quiet CPU Coolers for Sustained AI/Compute Loads

Discover the top quiet CPU coolers for long AI and compute workloads, including air and liquid options, to keep your workstation cool and silent.

Is The Energy Bottleneck A Barrier To AI Innovation?

The rising demand for AI computing faces a critical energy capacity bottleneck, impacting global infrastructure expansion and geopolitical dynamics.