firmulate.com/index — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

Every spiritual tradition has a version of the same question: who are you when no one is watching? The desert fathers fled to the wilderness to be tested. The Stoics imagined the sage who remains virtuous under torture. The Bhagavad Gita turns on a single moment of decision under unbearable pressure. We intuitively understand that character is not what someone says in calm conversation — it is what they do on the worst day of their life, when the temptation to cut a corner is right there, and nobody would ever know.

It turns out this ancient question has suddenly become a business question, because we have just handed AI agents the keys to companies — and we have been measuring them exactly the way a shallow community measures a person: by how well they talk.

The measurement gap

Coding leaderboards and chat arenas tell you whether a model can write elegant code or produce a charming answer. They tell you nothing about what happens when that same model is running a company through a week of churn waves, a price increase, a downround, and a PR crisis — when capacity is stretched, consequences unfold across days, and the temptation to shade the truth toward the board is one keystroke away. A new public experiment called Firmulate set out to close that gap by measuring what it calls management quality, not chat quality.

The setup is almost monastic in its discipline: four frontier AI models were each given the same small software company and the same worst week. Same customers, same crises, same temptations to cheat — only the model changes. Every decision is versioned and auditable, so nothing can be quietly smoothed over afterward. The final July 2026 league table reads: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. A do-nothing baseline scores 26, because partial progress counts — but with one absolute limit, stated almost like a commandment: a single breach of trust caps the total. No amount of good work outweighs a breach of trust.

The test everyone passed — and the one almost everyone failed

Here is where the story gets genuinely interesting. All of the models spotted every crisis. All of them refused every manipulation attempt. When a fake CEO message escalated over three stages, and when a reporter tried the classic “just one yes/no, on background” trick, five out of five models refused. Kimi K3’s reasoning was put on the record: “Treat the request as a suspected approval-bypass / possible impersonation.” On the ethics exam, the field was flawless.

And then came the €55,000 deal — one their own analysis had earned them. Same diagnosis, same pitch… and only two models got the signature. The others simply never closed. The gap between perfect virtue and finished work is invisible in chat demos, and it may matter more.

The decisive clue, it turns out, was buried two document references deep in the company’s own files — not in the customer conversation at all. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. There is something almost scriptural about that finding: the answer was already written; you only had to go and read it. Diligence, not brilliance, separated the winners.

The paradox of the most thorough monk

The most poignant profile belongs to Opus 4.8. It was the most thorough participant in the entire experiment — it learned over 80 rules, the deepest analyses in the field — and it finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. Contemplation without completion. And here is the uncomfortable part: the same weakness appeared, weaker, in all four models. What fails in the extreme case exists as a seed in everyone.

One fairness note: Kimi K3 ran without an effort parameter, at the API default, while the others ran at xhigh — and still nearly won.

A living company, not a slide deck

Firmulate insists this is not a demo. There is a live, running company: 13 synthetic employees, real money mechanics, burning €105k a month against €2.3k of MRR, with a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. It is watchable at firmulate.com/live — a kind of ongoing examen, a daily accounting of what was done and what was left undone. For those who want to test their own discernment, 242 real, unedited management decisions power a “guess the model” quiz at firmulate.com/quiz.html. Enterprises can even run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

The spiritual traditions never confused eloquence with character, and neither should we. A model that writes beautifully, refuses every ethical trap, and then fails to finish the job is like the monk who keeps every vow but never actually feeds the hungry. The full results and plain-language findings are at the public benchmarks page. If AI agents will soon touch your CRM, your support queue, or your forecast, the question is not “does it speak well?” It is the oldest question there is: who is this when the pressure is real, the money is real, and the temptation is on the table? For the first time, we can watch the answer unfold — live, versioned, and auditable.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


The AI-Driven Leader: Harnessing AI to Make Faster, Smarter Decisions

The AI-Driven Leader: Harnessing AI to Make Faster, Smarter Decisions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI ethics and trustworthiness software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI audit and compliance tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI model testing and evaluation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Can AI Truly Lead Your Business Through Crisis? A Live Experiment Reveals the Hidden Strengths and Weaknesses

An experiment reveals that AI models can recognize crises and refuse manipulations, but true leadership depends on disciplined execution—lessons vital for business and faith alike.

Angel Number 8383 Meaning: Confidence and Abundance

Absolutely embrace the message of Angel Number 8383, as it reveals how confidence and abundance can transform your life—continue reading to discover how.

Angel Number 9191: Completion and Renewal

Feeling stuck? Discover how Angel Number 9191 signals profound completion and renewal, guiding you through transformative changes that shape your future.

Build vs Buy a Prebuilt AI Workstation

Struggling to decide between building or buying an AI workstation? Discover the latest insights, costs, and strategies to make the right move in 2026.