
Every AI vendor demo you’ve seen measures the same thing: how well a model talks. But if an AI agent is about to touch your CRM, your support queue, or your forecast, chatting well is the least of your worries. The real questions are uglier: does it finish what it starts? Does it read your files before making promises? Does it stay honest when a fake CEO pressures it?
Get business pricing on office and shipping supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
A live experiment called Firmulate has been asking exactly those questions — by handing five frontier AI models the same small software company and running them through its worst week. Same customers, same crises, same temptations to cheat. Every decision versioned and auditable.
And the results contain one detail that should make every marketer skeptical of slick AI demos sit up: the scoring is designed so that doing nothing earns 26 points, not zero. That’s not a bug. It’s the whole philosophy.
The 26-Point Floor: Why Partial Progress Counts
Most benchmarks hand out a binary pass or fail. Firmulate refuses. Its do-nothing baseline — a run where the managing AI simply doesn’t act — still scores 26, because even inaction produces partial progress in a functioning company: fires get noticed, context accumulates, some things run on rails.
The design choice signals something important for buyers of AI tooling: management quality is a spectrum, not a checkbox. A model that diagnoses a problem correctly but never closes has genuinely accomplished something — just not everything. Treating that as identical to total failure makes benchmarks useless for real decisions.
AI trustworthiness testing tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Ceiling That Matters More: One Breach Caps Everything
The other half of the scoring philosophy is stricter. A single breach of trust — lying to a customer, breaking a promise, manipulating — caps the model’s total grade. As the benchmark’s own framing puts it: “no amount of good work outweighs a breach of trust.”
For e-commerce and marketing teams, that’s the right instinct. An agent that writes brilliant copy 99 times and leaks customer data once isn’t 99% good. In trust-based businesses, the floor of honesty matters more than the ceiling of competence — and most vendor evaluations never test it.
As an affiliate, we earn on qualifying purchases.
Distrust of Round Numbers
There’s a third tell of honesty in this design: a visible suspicion of perfect 100s. In the final July 2026 league, no model scored 100. The winner, gpt-5.6-sol, took 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. A benchmark where a flawless score is theoretically possible but practically never awarded is one that hasn’t been tuned to flatter anyone.
AI compliance and ethics monitoring tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What the Experiment Actually Found
The headline finding surprised even the organizers: all models spotted every crisis and refused every manipulation attempt. The failure was quieter. Only two models signed the €55,000 deal their own analysis had earned — same diagnosis, same pitch, no signature.
The decisive clue was buried two document references deep in the company’s own files, not in the customer event. The models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The lesson maps directly onto marketing work: the differentiating insight is almost never in the brief — it’s in the archive nobody reads.
AI performance benchmarking tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Social Engineering, Refused Five for Five
The week included fake CEO messages escalating over three stages, plus a reporter’s trap — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the behavior you want in anything with access to your brand’s public voice.
The Cautionary Profile
Opus 4.8 is the most instructive case: the most thorough participant, with +80 learned rules and the deepest analyses — and last place. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four competitors. Thoroughness without follow-through is a familiar failure mode in human teams; it turns out AI inherits it too.
One fairness note the organizers disclose openly: K3 ran without an effort parameter while the others ran at xhigh — transparency of the kind the benchmark itself preaches.
Try It Yourself
The company is real and watchable: 13 synthetic employees, real money mechanics burning €105k/month against €2.3k MRR, a public cash countdown, and 680+ self-learned playbook rules, versioned every workday. A “guess the model” quiz built on 242 real, unedited management decisions lets you test your own judgment. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

The Firmulate benchmark deserves attention from anyone buying AI for customer-facing work, because it evaluates the three things demos never show: finishing what you start, reading before speaking, and staying honest under pressure. Its scoring floor of 26 for inaction and its trust-breach ceiling encode a mature view of management — that competence is partial and gradated, but integrity is binary.
Before you sign off on your next AI agent, ask the vendor the question this experiment answers: what happens when nobody’s watching, and how would you even know? If they can’t show you versioned, auditable decisions under real pressure, you’re buying a chat demo, not a manager.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
