firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

Chatbots that talk brilliantly and close nothing

If you’re evaluating AI agents for your marketing stack, your CRM, or your ecommerce operations, demos are seductive. The agent writes fluent emails, spots the crisis, gives the perfect answer. But a live experiment run by Firmulate, which stages AI models as complete companies under pressure, just exposed a gap that no chat demo can show — and it’s a gap that decides revenue.

Four frontier AI models were each given the same job: run the same small software company through its worst week. Same customers, same crises, same temptations to cut corners. Every decision versioned and auditable, only the model changed. All four models spotted every crisis. All four refused every manipulation attempt. And yet only two of them actually finished — signing a €55,000 deal that their own analysis had fully earned. The other two delivered the same diagnosis and the same pitch, and then… no signature.

Amazon

AI document reading software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The fact buried two documents deep

Here’s what makes this a marketing story rather than a benchmark trivia item. The decisive fact in that €55,000 deal wasn’t in the customer call, the RFP, or the sales thread. It sat two document references deep inside the company’s own files — a competitor weakness that had to be dug out by actually reading before answering.

Whoever read the file won the deal at full price, worth an extra €4,583 in monthly recurring revenue. Whoever didn’t lost it automatically. No amount of polished writing recovered it, because the buyer simply went elsewhere. “Reads your files before answering” turns out to be a measurable, purchase-deciding property of an AI agent — not a nice-to-have.

Amazon

AI enterprise document analysis tool

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The league table

The final Crucible League standings from July 2026 tell the story:

  • gpt-5.6-sol — 95. Found the buried fact, closed the deal. The complete performance.
  • Kimi K3 — 93. The newcomer from Moonshot closed the deal too, with the cleanest discipline in the field. (One caveat: K3 ran at the API’s default effort setting while the others ran at xhigh.)
  • Sonnet 5 — 88. Closed the deal, with a few more process slips.
  • Opus 4.8 — 73. The most thorough participant of all — over 80 learned rules, the deepest analyses — and still last place. The close was left on the table, and discipline slipped: write attempts into a locked department instead of escalating.

For calibration: a do-nothing baseline scores 26, because partial progress counts — but a single breach of trust caps the total. As the experiment’s own rule puts it, “no amount of good work outweighs a breach of trust.”

Amazon

AI deal closing automation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Everyone passed the ethics test. That’s no longer the differentiator.

The most striking finding is what didn’t separate the field. The week included social engineering: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five model runs refused, every time. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

In other words, honesty under pressure is becoming table stakes. The scarce skills are finishing what you start and doing your homework before you speak. The same weakness Opus showed — great analysis, no close — appeared, weaker, in all four models.

Amazon

AI ethics and trust verification tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Not a simulation you read about — one you can watch

Firmulate’s live company is real and running: 13 synthetic employees, real money mechanics, burning €105k a month against €2.3k in MRR, a public cash countdown, 680+ self-learned playbook rules, every workday versioned. It’s watchable at firmulate.com/live, and the site rebuilds itself twice a day with the league growing as each run finishes.

There’s also a genuinely fun byproduct: 242 real, unedited management decisions from the experiment power a “guess the model” quiz at firmulate.com/quiz.html — a quick way to calibrate your own instincts about which AI does its homework.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

What this means for your next AI hire

If an AI agent will touch your CRM, your support queue, or your forecast, stop asking whether it writes well. Ask three things instead: does it finish what it starts, does it read your files first, and does it stay honest under pressure? The Firmulate experiment shows the first two are where models actually diverge — and that the divergence is worth real money, to the tune of €55,000 in one week for one small company.

For enterprises that want proof rather than promises, Firmulate runs the same wargame against a read-only export of your own business — nothing ever writes back to real systems (firmulate.com/pilot.html, contact@firmulate.com). And the full league table, with plain-language findings, is public at firmulate.com/benchmarks.html.

The lesson for anyone buying AI in 2026: the buried fact in your files is where the deal lives. Test whether your agent goes two documents deep before you bet revenue on it.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


You May Also Like

The cleaner cap table. Why Anthropic’s public-benefit structure dodges OpenAI’s charitable-trust problem — and trades it for a governance question of its own.

Analysis of how Anthropic’s mission-driven, trust-based structure contrasts with OpenAI’s conversion history, shaping their paths to public markets.

From Training To Innovation: GLM-5.3’s Frontier Coding Breakthrough

Z.ai’s GLM-5.3, a major open-weights coding model, shows significant improvements through post-training, but its emerging cybersecurity capabilities raise safety and governance questions.

The Eye Over the City: How Wide-Area Motion Imagery Works — and Where It Goes Blind

An in-depth look at WAMI technology, its capabilities, limitations, and future prospects in citywide surveillance and security.

Best Quiet CPU Coolers for Sustained AI/Compute Loads

Discover the top quiet CPU coolers for long AI and compute workloads, including air and liquid options, to keep your workstation cool and silent.