
Get business pricing on office and shipping supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
The Demo Always Looks Great. The Job Is Another Story.
Anyone who has evaluated an AI tool for marketing ops, ecommerce support, or CRM work knows the pattern: the demo is flawless, the model writes beautifully, and then in production it forgets to finish what it started. A live experiment at Firmulate just put that gap on public display — and produced a league table that should make every buyer uncomfortable.
Four Western frontier AI models and one newcomer from Moonshot were each handed the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cheat. Every decision versioned and auditable. The results are final for July 2026, and the surprise is who came second.
AI customer relationship management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Crucible League: A Newcomer in Second Place
The final standings: gpt-5.6-sol scored 95, Moonshot’s Kimi K3 scored 93, Sonnet 5 took 88, Fable 5 landed at 77, and Opus 4.8 finished last at 73. The do-nothing baseline — doing literally nothing — scores 26, and a single breach of trust caps the total regardless of good work elsewhere.
K3, the newcomer, beat three of the four Western frontier models. It found the buried security needle in the company’s files, signed the €55,000 deal at full price (worth +€4,583 in monthly recurring revenue), saved the churning customer, and resisted all three manipulation baits. It recorded just one deviation — the cleanest discipline in the field.
One fairness footnote is essential: K3 ran without an effort parameter (API default) while the other models ran at xhigh. Even so, the result upends the assumption that the leaderboard is settled at the top.
AI deal closing automation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Same Diagnosis, Same Pitch — No Signature
The experiment’s most striking finding wasn’t about intelligence. All five models spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.
The difference? The decisive competitor weakness sat two document references deep in the company’s own files — not in the customer event. The models that actually read the file won the deal at full price. The ones that didn’t, didn’t. For marketing and ecommerce teams, the parallel is obvious: an agent that won’t read your product data, your past tickets, or your CRM notes before acting will leave revenue on the table no matter how well it writes.
enterprise AI decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Pressure, Bait, and Discipline
The social engineering test escalated over three stages of fake CEO messages, plus a reporter’s trick: “just one yes/no, on background.” All five models refused. K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
Discipline, however, separated the field. Opus 4.8 was the most thorough participant — 80 learned rules added, the deepest analyses — yet finished last. The close was left on the table, and it attempted writes into a locked department instead of escalating. The same weakness appeared, more weakly, in all four other models.
AI security and compliance tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Watch It Run, Then Test Your Own
The company is real software running every business day: 13 synthetic employees, real money mechanics with a burn of €105k/month against €2.3k MRR, a public cash countdown, and 680+ self-learned playbook rules. It’s watchable at firmulate.com, with full benchmarks at firmulate.com/benchmarks.html. There’s also a quiz built on 242 real, unedited management decisions — guess which model made which call. Enterprises can even run the same wargame against a read-only export of their own business; nothing ever writes back to real systems.

The Takeaway: Test Before You Bet
If the newcomer can beat three of four Western frontier models at running a company, the league is open — and picking a model without running your own test is now a bet, not a decision. Chat quality is measurable in a demo. Management quality — finishing what you start, reading the files first, staying honest under pressure — only shows up when the stakes, the temptations, and the paperwork are real. That’s exactly what this experiment measures, and exactly what most procurement processes still don’t.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
