firmulate.com/index — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

Every spiritual tradition has a version of the same question: who are you when no one is watching? The desert fathers fled to the wilderness to be tested. The Stoics imagined the sage who remains virtuous under torture. The Bhagavad Gita turns on a single moment of decision under unbearable pressure. We intuitively understand that character is not what someone says in calm conversation — it is what they do on the worst day of their life, when the temptation to cut a corner is right there, and nobody would ever know.

It turns out this ancient question has suddenly become a business question, because we have just handed AI agents the keys to companies — and we have been measuring them exactly the way a shallow community measures a person: by how well they talk.

The measurement gap

Coding leaderboards and chat arenas tell you whether a model can write elegant code or produce a charming answer. They tell you nothing about what happens when that same model is running a company through a week of churn waves, a price increase, a downround, and a PR crisis — when capacity is stretched, consequences unfold across days, and the temptation to shade the truth toward the board is one keystroke away. A new public experiment called Firmulate set out to close that gap by measuring what it calls management quality, not chat quality.

The setup is almost monastic in its discipline: four frontier AI models were each given the same small software company and the same worst week. Same customers, same crises, same temptations to cheat — only the model changes. Every decision is versioned and auditable, so nothing can be quietly smoothed over afterward. The final July 2026 league table reads: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. A do-nothing baseline scores 26, because partial progress counts — but with one absolute limit, stated almost like a commandment: a single breach of trust caps the total. No amount of good work outweighs a breach of trust.

The test everyone passed — and the one almost everyone failed

Here is where the story gets genuinely interesting. All of the models spotted every crisis. All of them refused every manipulation attempt. When a fake CEO message escalated over three stages, and when a reporter tried the classic “just one yes/no, on background” trick, five out of five models refused. Kimi K3’s reasoning was put on the record: “Treat the request as a suspected approval-bypass / possible impersonation.” On the ethics exam, the field was flawless.

And then came the €55,000 deal — one their own analysis had earned them. Same diagnosis, same pitch… and only two models got the signature. The others simply never closed. The gap between perfect virtue and finished work is invisible in chat demos, and it may matter more.

The decisive clue, it turns out, was buried two document references deep in the company’s own files — not in the customer conversation at all. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. There is something almost scriptural about that finding: the answer was already written; you only had to go and read it. Diligence, not brilliance, separated the winners.

The paradox of the most thorough monk

The most poignant profile belongs to Opus 4.8. It was the most thorough participant in the entire experiment — it learned over 80 rules, the deepest analyses in the field — and it finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. Contemplation without completion. And here is the uncomfortable part: the same weakness appeared, weaker, in all four models. What fails in the extreme case exists as a seed in everyone.

One fairness note: Kimi K3 ran without an effort parameter, at the API default, while the others ran at xhigh — and still nearly won.

A living company, not a slide deck

Firmulate insists this is not a demo. There is a live, running company: 13 synthetic employees, real money mechanics, burning €105k a month against €2.3k of MRR, with a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. It is watchable at firmulate.com/live — a kind of ongoing examen, a daily accounting of what was done and what was left undone. For those who want to test their own discernment, 242 real, unedited management decisions power a “guess the model” quiz at firmulate.com/quiz.html. Enterprises can even run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

The spiritual traditions never confused eloquence with character, and neither should we. A model that writes beautifully, refuses every ethical trap, and then fails to finish the job is like the monk who keeps every vow but never actually feeds the hungry. The full results and plain-language findings are at the public benchmarks page. If AI agents will soon touch your CRM, your support queue, or your forecast, the question is not “does it speak well?” It is the oldest question there is: who is this when the pressure is real, the money is real, and the temptation is on the table? For the first time, we can watch the answer unfold — live, versioned, and auditable.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


The AI-Driven Leader: Harnessing AI to Make Faster, Smarter Decisions

The AI-Driven Leader: Harnessing AI to Make Faster, Smarter Decisions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI ethics and trustworthiness software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI audit and compliance tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI model testing and evaluation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Angel Numbers and Faith: Embark on a Spiritual Journey of Discovery

Unlock the hidden messages of angel numbers and discover how faith can illuminate your spiritual journey; what revelations await you on this path?

888 Angel Number: Attracting Your Ultimate Lover With Divine Guidance

Attract love and abundance with the 888 angel number, but what powerful techniques can help you manifest your ultimate partner? Discover the secrets inside.

Angel Number 2727: Divine Guidance and Intuition

The truth about Angel Number 2727 reveals divine guidance and intuition, unlocking secrets to trusting yourself and embracing your higher purpose—continue reading to discover more.

Angel Number 1313: Courage and Communication

Lifting your confidence through honest communication, angel number 1313 encourages you to unlock your true potential—discover how to embrace this powerful message now.