firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Every AI vendor demo you’ve seen measures the same thing: how well a model talks. But if an AI agent is about to touch your CRM, your support queue, or your forecast, chatting well is the least of your worries. The real questions are uglier: does it finish what it starts? Does it read your files before making promises? Does it stay honest when a fake CEO pressures it?

Buying for a business?Offer from Amazon

Get business pricing on office and shipping supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

A live experiment called Firmulate has been asking exactly those questions — by handing five frontier AI models the same small software company and running them through its worst week. Same customers, same crises, same temptations to cheat. Every decision versioned and auditable.

And the results contain one detail that should make every marketer skeptical of slick AI demos sit up: the scoring is designed so that doing nothing earns 26 points, not zero. That’s not a bug. It’s the whole philosophy.

The 26-Point Floor: Why Partial Progress Counts

Most benchmarks hand out a binary pass or fail. Firmulate refuses. Its do-nothing baseline — a run where the managing AI simply doesn’t act — still scores 26, because even inaction produces partial progress in a functioning company: fires get noticed, context accumulates, some things run on rails.

The design choice signals something important for buyers of AI tooling: management quality is a spectrum, not a checkbox. A model that diagnoses a problem correctly but never closes has genuinely accomplished something — just not everything. Treating that as identical to total failure makes benchmarks useless for real decisions.

Amazon

AI trustworthiness testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Ceiling That Matters More: One Breach Caps Everything

The other half of the scoring philosophy is stricter. A single breach of trust — lying to a customer, breaking a promise, manipulating — caps the model’s total grade. As the benchmark’s own framing puts it: “no amount of good work outweighs a breach of trust.”

For e-commerce and marketing teams, that’s the right instinct. An agent that writes brilliant copy 99 times and leaks customer data once isn’t 99% good. In trust-based businesses, the floor of honesty matters more than the ceiling of competence — and most vendor evaluations never test it.

Amazon

AI model evaluation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Distrust of Round Numbers

There’s a third tell of honesty in this design: a visible suspicion of perfect 100s. In the final July 2026 league, no model scored 100. The winner, gpt-5.6-sol, took 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. A benchmark where a flawless score is theoretically possible but practically never awarded is one that hasn’t been tuned to flatter anyone.

Amazon

AI compliance and ethics monitoring tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Experiment Actually Found

The headline finding surprised even the organizers: all models spotted every crisis and refused every manipulation attempt. The failure was quieter. Only two models signed the €55,000 deal their own analysis had earned — same diagnosis, same pitch, no signature.

The decisive clue was buried two document references deep in the company’s own files, not in the customer event. The models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The lesson maps directly onto marketing work: the differentiating insight is almost never in the brief — it’s in the archive nobody reads.

Amazon

AI performance benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Social Engineering, Refused Five for Five

The week included fake CEO messages escalating over three stages, plus a reporter’s trap — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the behavior you want in anything with access to your brand’s public voice.

The Cautionary Profile

Opus 4.8 is the most instructive case: the most thorough participant, with +80 learned rules and the deepest analyses — and last place. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four competitors. Thoroughness without follow-through is a familiar failure mode in human teams; it turns out AI inherits it too.

One fairness note the organizers disclose openly: K3 ran without an effort parameter while the others ran at xhigh — transparency of the kind the benchmark itself preaches.

Try It Yourself

The company is real and watchable: 13 synthetic employees, real money mechanics burning €105k/month against €2.3k MRR, a public cash countdown, and 680+ self-learned playbook rules, versioned every workday. A “guess the model” quiz built on 242 real, unedited management decisions lets you test your own judgment. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The Firmulate benchmark deserves attention from anyone buying AI for customer-facing work, because it evaluates the three things demos never show: finishing what you start, reading before speaking, and staying honest under pressure. Its scoring floor of 26 for inaction and its trust-breach ceiling encode a mature view of management — that competence is partial and gradated, but integrity is binary.

Before you sign off on your next AI agent, ask the vendor the question this experiment answers: what happens when nobody’s watching, and how would you even know? If they can’t show you versioned, auditable decisions under real pressure, you’re buying a chat demo, not a manager.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Significance Of OpenAI Securing The Fields Medal Winner In AI

OpenAI has reportedly recruited the latest Fields Medal-winning mathematician, signaling a focus on advanced reasoning. ByteDance’s new scientist program highlights US-China AI talent competition.

Waves, Not a Wall: Inside DeepMind’s Map From AGI to Superintelligence

DeepMind researchers present a framework mapping the transition from AGI to superintelligence, highlighting pathways, challenges, and uncertainties.

The Anthropic-Blackstone-Goldman JV: Reverse-Engineering the $1.5B Enterprise AI Services Structure

Anthropic, Blackstone, Hellman & Friedman, and Goldman Sachs launched a $1.5 billion AI enterprise services company, embedding Anthropic engineers inside a new standalone entity.

Inside China’s Fast-Paced AI Model Deployment: Signal’s Four Open Versions

Chinese labs released four frontier-class open-weight AI models in eight weeks, transforming the global AI landscape and impacting deployment strategies.