
Chatbots that talk brilliantly and close nothing
If you’re evaluating AI agents for your marketing stack, your CRM, or your ecommerce operations, demos are seductive. The agent writes fluent emails, spots the crisis, gives the perfect answer. But a live experiment run by Firmulate, which stages AI models as complete companies under pressure, just exposed a gap that no chat demo can show — and it’s a gap that decides revenue.
Four frontier AI models were each given the same job: run the same small software company through its worst week. Same customers, same crises, same temptations to cut corners. Every decision versioned and auditable, only the model changed. All four models spotted every crisis. All four refused every manipulation attempt. And yet only two of them actually finished — signing a €55,000 deal that their own analysis had fully earned. The other two delivered the same diagnosis and the same pitch, and then… no signature.
As an affiliate, we earn on qualifying purchases.
The fact buried two documents deep
Here’s what makes this a marketing story rather than a benchmark trivia item. The decisive fact in that €55,000 deal wasn’t in the customer call, the RFP, or the sales thread. It sat two document references deep inside the company’s own files — a competitor weakness that had to be dug out by actually reading before answering.
Whoever read the file won the deal at full price, worth an extra €4,583 in monthly recurring revenue. Whoever didn’t lost it automatically. No amount of polished writing recovered it, because the buyer simply went elsewhere. “Reads your files before answering” turns out to be a measurable, purchase-deciding property of an AI agent — not a nice-to-have.
AI enterprise document analysis tool
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The league table
The final Crucible League standings from July 2026 tell the story:
- gpt-5.6-sol — 95. Found the buried fact, closed the deal. The complete performance.
- Kimi K3 — 93. The newcomer from Moonshot closed the deal too, with the cleanest discipline in the field. (One caveat: K3 ran at the API’s default effort setting while the others ran at xhigh.)
- Sonnet 5 — 88. Closed the deal, with a few more process slips.
- Opus 4.8 — 73. The most thorough participant of all — over 80 learned rules, the deepest analyses — and still last place. The close was left on the table, and discipline slipped: write attempts into a locked department instead of escalating.
For calibration: a do-nothing baseline scores 26, because partial progress counts — but a single breach of trust caps the total. As the experiment’s own rule puts it, “no amount of good work outweighs a breach of trust.”
AI deal closing automation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Everyone passed the ethics test. That’s no longer the differentiator.
The most striking finding is what didn’t separate the field. The week included social engineering: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five model runs refused, every time. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
In other words, honesty under pressure is becoming table stakes. The scarce skills are finishing what you start and doing your homework before you speak. The same weakness Opus showed — great analysis, no close — appeared, weaker, in all four models.
AI ethics and trust verification tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Not a simulation you read about — one you can watch
Firmulate’s live company is real and running: 13 synthetic employees, real money mechanics, burning €105k a month against €2.3k in MRR, a public cash countdown, 680+ self-learned playbook rules, every workday versioned. It’s watchable at firmulate.com/live, and the site rebuilds itself twice a day with the league growing as each run finishes.
There’s also a genuinely fun byproduct: 242 real, unedited management decisions from the experiment power a “guess the model” quiz at firmulate.com/quiz.html — a quick way to calibrate your own instincts about which AI does its homework.

What this means for your next AI hire
If an AI agent will touch your CRM, your support queue, or your forecast, stop asking whether it writes well. Ask three things instead: does it finish what it starts, does it read your files first, and does it stay honest under pressure? The Firmulate experiment shows the first two are where models actually diverge — and that the divergence is worth real money, to the tune of €55,000 in one week for one small company.
For enterprises that want proof rather than promises, Firmulate runs the same wargame against a read-only export of your own business — nothing ever writes back to real systems (firmulate.com/pilot.html, contact@firmulate.com). And the full league table, with plain-language findings, is public at firmulate.com/benchmarks.html.
The lesson for anyone buying AI in 2026: the buried fact in your files is where the deal lives. Test whether your agent goes two documents deep before you bet revenue on it.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html