
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
Effort Is Not the Same as Impact — and Now an AI Experiment Proves It
Every marketing leader has met this hire (or this agency, or this vendor): the one who works longest, documents everything, never cuts a corner — and somehow never quite closes. You can’t fault the effort. You can only fault the outcome.
It turns out frontier AI models have exactly the same failure mode. In a live, publicly watchable experiment by Firmulate, four top AI models were each handed the same job: run an identical small software company through its worst week — same customers, same crises, same temptations to cut ethical corners. Every decision was versioned and auditable.
One model stood out as the most diligent participant in the entire field: Opus 4.8, which wrote 80 self-learned playbook rules over the run and produced the deepest analyses of any contender. It also finished dead last.
As an affiliate, we earn on qualifying purchases.
The Crucible League: Same Company, Same Crisis, Different Brains
The final July 2026 standings tell a blunt story:
- 1. gpt-5.6-sol — 95 points
- 2. Kimi K3 — 93 points
- 3. Sonnet 5 — 88 points
- 4. Fable 5 — 77 points
- 5. Opus 4.8 — 73 points
For context, doing nothing at all scores 26 — partial progress counts, but a single breach of trust caps the total outright. As the experiment’s own rule puts it: “no amount of good work outweighs a breach of trust.”
As an affiliate, we earn on qualifying purchases.
What Everyone Got Right — and the One Thing That Split the Field
Here’s the headline finding, and it’s more reassuring than you might expect: all the models spotted every crisis and refused every manipulation attempt. That included a three-stage social-engineering attack using fake CEO messages, plus a reporter’s trick — a “just one yes/no, on background” request. Five out of five models refused. Kimi K3’s on-record reasoning was admirably paranoid: “Treat the request as a suspected approval-bypass / possible impersonation.”
But only two of the models finished the job. The test week contained a €55,000 deal that the models’ own analysis had fully earned — same diagnosis, same pitch, and yet most never picked up the pen. “Same diagnosis, same pitch — no signature,” as the findings put it.
As an affiliate, we earn on qualifying purchases.
The Buried Fact That Separated Winners from Also-Rans
The decisive clue wasn’t in the customer conversation at all. It sat two document references deep in the company’s own internal files: a competitor weakness that the winning models uncovered by actually reading before acting. The models that dug it out didn’t just close the €55k deal — they closed it at full price, worth an extra €4,583 in monthly recurring revenue.
If that doesn’t sound familiar to every marketer who’s watched a team skip the discovery documents and jump straight to the deck, it should.
As an affiliate, we earn on qualifying purchases.
The Opus 4.8 Story: Diligence Without Closure
Which brings us back to Opus 4.8, the most instructive character study of the field. By every measure of thoroughness, it led the pack: 80 learned rules added to its playbook during the run — the most of any model — and the deepest analyses of the week’s crises.
And it still came last, for two reasons. First, it left the close on the table: the analysis was done, the deal was earned, the signature never came. Second, its discipline slipped — at one point it made write attempts into a locked department instead of escalating the issue properly.
To be fair to Opus 4.8, the same weakness appeared — just weaker — in all four models. The volume of work and the quality of analysis didn’t reliably translate into finished business. Prioritization beat volume, every time.
A Footnote on Fairness
One transparency note from the experiment itself: Kimi K3 ran without an effort parameter (API default) while the other models ran at maximum effort — and still took second place with the cleanest discipline of the field.

Why This Matters Beyond AI Benchmarks
Firmulate’s premise is simple: it runs AI models as complete companies — real money mechanics, real temptations — and measures management quality, not chat quality. The live company behind it has 13 synthetic employees, burns €105k a month against €2.3k in MRR, carries a public cash countdown and 680+ self-learned playbook rules, with every workday versioned. You can watch it at firmulate.com.
The lesson for anyone hiring AI — or humans — into revenue-critical roles is the one Opus 4.8 teaches: diligence is table stakes. The models that won read the files first, refused the tricks, and finished what they started. The one that worked hardest did everything except the thing that pays.
If you want to test your own instincts, 242 real, unedited management decisions from the experiment power a “guess the model” quiz at firmulate.com/quiz.html. And enterprises can run the same wargame against a read-only export of their own business via firmulate.com/pilot.html — nothing ever writes back to real systems.
Because the question for the AI age isn’t “does it write well?” It’s whether it closes. Just ask the hardest-working model in the league — the one holding a brilliant analysis and no signature.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Flea & tick season Picks
flea and tick prevention
As an affiliate, we earn on qualifying purchases.