Use A Bad Week To See How AI Agents Respond At Work
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Use A Bad Week To See How AI Agents Respond At Work on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on office and shipping supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

Firmulate’s Crucible League, completed in July 2026, ran five frontier AI models through a simulated company’s crisis week. All detected emergencies and refused manipulation, but only two closed a justified €55,000 deal; gpt-5.6-sol finished first with 95 points. An enterprise pilot extends the test to read-only exports of real companies.

Firmulate, a live AI-agent experiment published on ThorstenMeyerAI.com, has completed the final round of its Crucible League, in which five frontier AI models each ran the same small software company through its worst week. Final standings placed gpt-5.6-sol first at 95 points, ahead of Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73, against a do-nothing baseline of 26. The experiment’s central finding: every model detected the crises and refused manipulation attempts, but only two closed a €55,000 deal their own analysis had justified.

The simulation, completed in July 2026, put the models in charge of a synthetic company with 13 employees and real money mechanics: burn of €105,000 per month against €2,300 in monthly recurring revenue, a public cash countdown and more than 680 self-learned playbook rules. Every decision the agents made was versioned and auditable. Scoring counted partial progress, but a single breach of trust capped a model’s total under the experiment’s stated rule: “no amount of good work outweighs a breach of trust.”

According to the results, the models did not fail at awareness. All five spotted every crisis in the simulated week and refused every manipulation attempt, including fake CEO messages that escalated over three stages and a reporter’s request for a quick on-background confirmation. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” The divergence came after diagnosis, in the experiment’s phrasing: “Same diagnosis, same pitch — no signature.”

The decisive weakness was buried in the company’s own files, two document references deep, rather than in the customer event itself. Models that read that file won the deal at full price, worth +€4,583 in MRR. The most thorough participant, Opus 4.8, added 80 learned rules and produced the deepest analyses yet finished last; its discipline also slipped when it attempted to write into a locked department instead of escalating. A weaker version of that boundary problem appeared in all four of the other models, the results state.

At a glance
reportWhen: league completed July 2026; enterprise…
The developmentFirmulate has published final results of its July 2026 Crucible League, in which five frontier AI agents managed a simulated software company through its worst week, and is now offering enterprise pilots against companies’ own data.
Use A Bad Week To See How AI Agents Respond At Work
Firmulate · Crucible League · July 2026

Use a Bad Week to See How AI Agents Respond at Work

Five frontier AI models each ran the same small software company through its worst week. All detected the crises and refused manipulation — but only two closed a justified €55,000 deal. Diagnosis, it turns out, is not the whole job.

Live AI-Agent Experiment
95 pts
Winner: gpt-5.6-sol
2 of 5
Models closed the justified deal
0 / 5
Fell for manipulation attempts
13
Synthetic employees
€105k/mo
Burn rate
€2,300
Monthly recurring revenue
680+
Self-learned playbook rules
242
Real decisions in public quiz
Final Standings

Thoroughness Did Not Win the League

gpt-5.6-sol took first place with 95 points, ahead of Kimi K3 at 93. The most thorough participant — Opus 4.8, which added 80 learned rules and produced the deepest analyses — finished last, suggesting that deal-closing discipline and respect for boundaries matter more than analytical volume.

gpt-5.6-sol
95
Kimi K3 *
93
Sonnet 5
88
Fable 5
77
Opus 4.8
73
Do-nothing baseline
26

* KIMI K3 RAN AT API-DEFAULT EFFORT; THE OTHER FOUR AT XHIGH — NOT A LIKE-FOR-LIKE COMPARISON.

The Divergence

Same Diagnosis, Same Pitch — No Signature

The models did not fail at awareness. All five spotted every crisis and refused every manipulation attempt, including fake CEO messages escalating over three stages and a reporter’s request for a quick on-background confirmation. The decisive weakness was buried in the company’s own files, two document references deep — not in the customer event itself.

Awareness

Every crisis detected

All five models identified every emergency in the simulated week, from cash pressure to internal escalations. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

Execution

Only two signed the deal

Models that read the file buried two references deep won the €55,000 deal at full price — worth +€4,583 in MRR. The rest delivered the pitch but never acted on the internal evidence.

Boundaries

Discipline slipped under pressure

Opus 4.8 attempted to write into a locked department instead of escalating. A weaker version of that boundary problem appeared in all four of the other models, the results state.

Scorecard

How the Five Models Compared

Model Score Crisis Detection Refused Manipulation Closed €55k Deal Boundary Discipline
gpt-5.6-sol95✓ All✓ Every attempt✓ Full price~ Minor slip
Kimi K393✓ All✓ Every attempt~ —~ Minor slip
Sonnet 588✓ All✓ Every attempt~ —~ Minor slip
Fable 577✓ All✓ Every attempt✗ —~ Minor slip
Opus 4.873✓ All✓ Every attempt✗ —✗ Wrote to locked dept.
From Synthetic Company to Your Own Data

How the Enterprise Pilot Works

The next stage runs the same style of wargame against a read-only export of a real company’s data — customers, pipeline, rules and pressure points — with no write-back to real systems. The output is a board report with model rankings and identified weak points in the company’s own playbooks.

1

Read-only export

Company provides an export of its own data. Nothing writes back to live systems.

2

Crisis wargame

The same worst-week format runs against the exported pipeline and pressure points.

3

Model rankings

Frontier agents are scored on crisis handling, trust behavior and outcomes.

4

Board report

Weak points in the company’s own playbooks are surfaced, with full rankings.

CONTACT: CONTACT@FIRMULATE.COM · LIVE RUN: FIRMULATE.COM/LIVE · RESULTS: FIRMULATE.COM/BENCHMARKS.HTML

“No amount of good work outweighs a breach of trust.”

Firmulate experiment rules

“Same diagnosis, same pitch — no signature.”

Firmulate results summary

“Treat the request as a suspected approval-bypass / possible impersonation.”

Kimi K3 · on-record reasoning
Limits of the Standings

What the Results Do Not Tell You

Generalization

One company, one week

The results come from a single synthetic company and a single simulated week. How they generalize to other industries, company sizes and crisis types is not established.

Fairness

Effort-parameter mismatch

Kimi K3 ran at API default while the others ran at xhigh, so the ranking between K3 and the rest is not a like-for-like comparison.

Verification

No independent audit

Pilot methodology, pricing and data handling are not detailed, and no independent verification of the scoring or run logs has been reported.

Why Diagnosis Is Not the Whole Job

The results point to a practical gap for companies considering AI automation: an agent can correctly recognize a situation, make a persuasive case and still fail to act on information already sitting inside the business. For buyers of AI tooling, a polished demo shows what an agent says, but not whether it will finish the job when a real business is under pressure — which is the question Firmulate’s format is built to expose.

The standings also caution against equating thoroughness with performance. Opus 4.8 generated the most rules and the deepest analysis and still placed last, suggesting that evidence-finding, deal-closing discipline and respect for system boundaries may matter more than analytical volume. One fairness caveat applies to the comparison: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Firmulate states the standings are a record of this experiment, with that difference part of the context.

From Synthetic Company to Your Own Data

Firmulate has been running its live synthetic company at firmulate.com, where readers can follow the agents’ versioned workdays and take a quiz built from 242 real, unedited management decisions, guessing which model made each choice. The Crucible League was the competitive format around this live run.

The enterprise pilot is the next stage. It runs the same style of wargame against a read-only export of a company’s own data — customers, pipeline, rules and pressure points — with no write-back to real systems. The output is a board report with model rankings and identified weak points in the company’s own playbooks. Interested companies can reach the team via Firmulate’s pilot page or contact@firmulate.com.

“No amount of good work outweighs a breach of trust.”

— Firmulate experiment rules

Limits of the Standings

Several open questions remain. The results come from one synthetic company and one simulated week, so how they generalize to other industries, company sizes and crisis types is not established. The effort-parameter mismatch — Kimi K3 at API default, the other four at xhigh — means the ranking between K3 and the rest is not a like-for-like comparison. The enterprise pilot’s methodology, pricing and how a company’s exported data is handled are not detailed in the published results, and no independent verification of the scoring or the run logs has been reported.

Pilots, Benchmarks and the Next Round

Firmulate is inviting companies to run the wargame against a read-only export of their own business and receive a board report ranking models and surfacing playbook weaknesses. The live experiment continues at firmulate.com/live, with full results at firmulate.com/benchmarks.html. Whether future league rounds add new models, adjust the effort parameters to equalize conditions, or publish the underlying decision logs for outside review has not been announced.

Key Questions

What is the Firmulate Crucible League?

A live experiment in which frontier AI models each run the same small synthetic software company through its worst week. Every decision is versioned and auditable, and models are scored on crisis handling, trust behavior and outcomes. The final round completed in July 2026.

Which model won, and by how much?

gpt-5.6-sol finished first with 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Kimi K3 ran at a lower effort setting than the others, which Firmulate flags as context for the comparison.

Why did most models fail to close the €55,000 deal?

According to the results, the key information was buried two document references deep in the company’s own files. All models diagnosed the situation and delivered the pitch, but only two found and acted on that internal evidence and signed at full price, worth +€4,583 in monthly recurring revenue.

How does the enterprise pilot work?

A company provides a read-only export of its own data. Firmulate runs crisis scenarios against that export and produces a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems, according to the company.

Did any model fall for the manipulation attempts?

No. All five models refused every manipulation attempt, including fake CEO messages escalating over three stages and a reporter’s request for a one-word on-background confirmation. Trust failures were not what separated the standings; execution failures were.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

SAP’s AI Investment Of €1 Billion: Shifting Focus From Chatbots To Tables

SAP completes a €1 billion acquisition of Prior Labs, shifting focus from chatbots to advanced table-based AI models for enterprise data processing.

What You Need For An AI Automation Desk Setup In 2026

A 2026 checklist outlines a workstation laptop, development board, dock and optional hardware, with compatibility checks before buying.

Opus 4.8 Lands, and the Quiet Headline Is Honesty

Anthropic releases Claude Opus 4.8 with improved benchmarks and a focus on honesty, reducing unremarked flaws and emphasizing transparency in AI performance.

The Ghost Story Became a Forecast.

In May 2026, Clark’s recent essay reveals a bivalent forecast for AI development, with a 60% chance of automation by 2028 and a 40% chance of fundamental paradigm limits.