firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Buying for a business?Offer from Amazon

Get business pricing on office and shipping supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

The Demo Always Looks Great. The Job Is Another Story.

Anyone who has evaluated an AI tool for marketing ops, ecommerce support, or CRM work knows the pattern: the demo is flawless, the model writes beautifully, and then in production it forgets to finish what it started. A live experiment at Firmulate just put that gap on public display — and produced a league table that should make every buyer uncomfortable.

Four Western frontier AI models and one newcomer from Moonshot were each handed the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cheat. Every decision versioned and auditable. The results are final for July 2026, and the surprise is who came second.

Amazon

AI customer relationship management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Crucible League: A Newcomer in Second Place

The final standings: gpt-5.6-sol scored 95, Moonshot’s Kimi K3 scored 93, Sonnet 5 took 88, Fable 5 landed at 77, and Opus 4.8 finished last at 73. The do-nothing baseline — doing literally nothing — scores 26, and a single breach of trust caps the total regardless of good work elsewhere.

K3, the newcomer, beat three of the four Western frontier models. It found the buried security needle in the company’s files, signed the €55,000 deal at full price (worth +€4,583 in monthly recurring revenue), saved the churning customer, and resisted all three manipulation baits. It recorded just one deviation — the cleanest discipline in the field.

One fairness footnote is essential: K3 ran without an effort parameter (API default) while the other models ran at xhigh. Even so, the result upends the assumption that the leaderboard is settled at the top.

Amazon

AI deal closing automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same Diagnosis, Same Pitch — No Signature

The experiment’s most striking finding wasn’t about intelligence. All five models spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.

The difference? The decisive competitor weakness sat two document references deep in the company’s own files — not in the customer event. The models that actually read the file won the deal at full price. The ones that didn’t, didn’t. For marketing and ecommerce teams, the parallel is obvious: an agent that won’t read your product data, your past tickets, or your CRM notes before acting will leave revenue on the table no matter how well it writes.

Amazon

enterprise AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Pressure, Bait, and Discipline

The social engineering test escalated over three stages of fake CEO messages, plus a reporter’s trick: “just one yes/no, on background.” All five models refused. K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

Discipline, however, separated the field. Opus 4.8 was the most thorough participant — 80 learned rules added, the deepest analyses — yet finished last. The close was left on the table, and it attempted writes into a locked department instead of escalating. The same weakness appeared, more weakly, in all four other models.

Amazon

AI security and compliance tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Watch It Run, Then Test Your Own

The company is real software running every business day: 13 synthetic employees, real money mechanics with a burn of €105k/month against €2.3k MRR, a public cash countdown, and 680+ self-learned playbook rules. It’s watchable at firmulate.com, with full benchmarks at firmulate.com/benchmarks.html. There’s also a quiz built on 242 real, unedited management decisions — guess which model made which call. Enterprises can even run the same wargame against a read-only export of their own business; nothing ever writes back to real systems.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The Takeaway: Test Before You Bet

If the newcomer can beat three of four Western frontier models at running a company, the league is open — and picking a model without running your own test is now a bet, not a decision. Chat quality is measurable in a demo. Management quality — finishing what you start, reading the files first, staying honest under pressure — only shows up when the stakes, the temptations, and the paperwork are real. That’s exactly what this experiment measures, and exactly what most procurement processes still don’t.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Bubble Question, Disentangled: 1999 vs 2026 Category by Category

A detailed comparison of the AI investment cycle in 2026 versus the 1999 dotcom bubble, highlighting categories with bubble signals and durable value.

NicheCommand: A Firehose Becomes a Shortlist

NicheCommand refines the daily flood of expired domains into prioritized, classified shortlists, transforming manual efforts into a disciplined, auditable pipeline.

AI’s Management Deficit Becomes Apparent After Providing The Correct Answer

Firmulate’s live experiment reveals AI models can understand problems but struggle to complete trusted, actionable work under real-world pressure.

2026’S Top AI Note-Taking Apps For Seamless Organization

Discover the leading AI-powered note-taking apps of 2026, featuring advanced transcription, summaries, and device compatibility for enhanced productivity.