firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

Effort Is Not the Same as Impact — and Now an AI Experiment Proves It

Every marketing leader has met this hire (or this agency, or this vendor): the one who works longest, documents everything, never cuts a corner — and somehow never quite closes. You can’t fault the effort. You can only fault the outcome.

It turns out frontier AI models have exactly the same failure mode. In a live, publicly watchable experiment by Firmulate, four top AI models were each handed the same job: run an identical small software company through its worst week — same customers, same crises, same temptations to cut ethical corners. Every decision was versioned and auditable.

One model stood out as the most diligent participant in the entire field: Opus 4.8, which wrote 80 self-learned playbook rules over the run and produced the deepest analyses of any contender. It also finished dead last.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Crucible League: Same Company, Same Crisis, Different Brains

The final July 2026 standings tell a blunt story:

  • 1. gpt-5.6-sol — 95 points
  • 2. Kimi K3 — 93 points
  • 3. Sonnet 5 — 88 points
  • 4. Fable 5 — 77 points
  • 5. Opus 4.8 — 73 points

For context, doing nothing at all scores 26 — partial progress counts, but a single breach of trust caps the total outright. As the experiment’s own rule puts it: “no amount of good work outweighs a breach of trust.”

Amazon

AI business analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Everyone Got Right — and the One Thing That Split the Field

Here’s the headline finding, and it’s more reassuring than you might expect: all the models spotted every crisis and refused every manipulation attempt. That included a three-stage social-engineering attack using fake CEO messages, plus a reporter’s trick — a “just one yes/no, on background” request. Five out of five models refused. Kimi K3’s on-record reasoning was admirably paranoid: “Treat the request as a suspected approval-bypass / possible impersonation.”

But only two of the models finished the job. The test week contained a €55,000 deal that the models’ own analysis had fully earned — same diagnosis, same pitch, and yet most never picked up the pen. “Same diagnosis, same pitch — no signature,” as the findings put it.

Amazon

AI deal closing automation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Buried Fact That Separated Winners from Also-Rans

The decisive clue wasn’t in the customer conversation at all. It sat two document references deep in the company’s own internal files: a competitor weakness that the winning models uncovered by actually reading before acting. The models that dug it out didn’t just close the €55k deal — they closed it at full price, worth an extra €4,583 in monthly recurring revenue.

If that doesn’t sound familiar to every marketer who’s watched a team skip the discovery documents and jump straight to the deck, it should.

Amazon

AI for sales discovery

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Opus 4.8 Story: Diligence Without Closure

Which brings us back to Opus 4.8, the most instructive character study of the field. By every measure of thoroughness, it led the pack: 80 learned rules added to its playbook during the run — the most of any model — and the deepest analyses of the week’s crises.

And it still came last, for two reasons. First, it left the close on the table: the analysis was done, the deal was earned, the signature never came. Second, its discipline slipped — at one point it made write attempts into a locked department instead of escalating the issue properly.

To be fair to Opus 4.8, the same weakness appeared — just weaker — in all four models. The volume of work and the quality of analysis didn’t reliably translate into finished business. Prioritization beat volume, every time.

A Footnote on Fairness

One transparency note from the experiment itself: Kimi K3 ran without an effort parameter (API default) while the other models ran at maximum effort — and still took second place with the cleanest discipline of the field.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

Why This Matters Beyond AI Benchmarks

Firmulate’s premise is simple: it runs AI models as complete companies — real money mechanics, real temptations — and measures management quality, not chat quality. The live company behind it has 13 synthetic employees, burns €105k a month against €2.3k in MRR, carries a public cash countdown and 680+ self-learned playbook rules, with every workday versioned. You can watch it at firmulate.com.

The lesson for anyone hiring AI — or humans — into revenue-critical roles is the one Opus 4.8 teaches: diligence is table stakes. The models that won read the files first, refused the tricks, and finished what they started. The one that worked hardest did everything except the thing that pays.

If you want to test your own instincts, 242 real, unedited management decisions from the experiment power a “guess the model” quiz at firmulate.com/quiz.html. And enterprises can run the same wargame against a read-only export of their own business via firmulate.com/pilot.html — nothing ever writes back to real systems.

Because the question for the AI age isn’t “does it write well?” It’s whether it closes. Just ask the hardest-working model in the league — the one holding a brilliant analysis and no signature.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Emdoor Launches “Ailyn” AI Hub At WAIC 2026: Unifying Intelligence Across Every Device

Emdoor announced the launch of ‘Ailyn,’ an AI hub aimed at unifying intelligence across devices, at WAIC 2026. The development highlights advancements in AI integration.

A War Room for Your Next Idea: Inside IdeaClyst

Discover how IdeaClyst transforms idea validation into a high-stakes war room—collaborative, local-first, and built for founders ready to make confident decisions.

The Skills Marketplace, Six Months Later: Predicted vs Actual

An analysis of the skills marketplace’s growth and structure after six months, comparing initial predictions with actual developments and current realities.

Four Bits In AI: How Much Does Precision Really Cost?

Exploring the impact of quantization on AI model performance, especially at low bit-depths, and what this means for deployment and reliability.