
Every spiritual tradition has a version of the same question: who are you when no one is watching? The desert fathers fled to the wilderness to be tested. The Stoics imagined the sage who remains virtuous under torture. The Bhagavad Gita turns on a single moment of decision under unbearable pressure. We intuitively understand that character is not what someone says in calm conversation — it is what they do on the worst day of their life, when the temptation to cut a corner is right there, and nobody would ever know.
It turns out this ancient question has suddenly become a business question, because we have just handed AI agents the keys to companies — and we have been measuring them exactly the way a shallow community measures a person: by how well they talk.
The measurement gap
Coding leaderboards and chat arenas tell you whether a model can write elegant code or produce a charming answer. They tell you nothing about what happens when that same model is running a company through a week of churn waves, a price increase, a downround, and a PR crisis — when capacity is stretched, consequences unfold across days, and the temptation to shade the truth toward the board is one keystroke away. A new public experiment called Firmulate set out to close that gap by measuring what it calls management quality, not chat quality.
The setup is almost monastic in its discipline: four frontier AI models were each given the same small software company and the same worst week. Same customers, same crises, same temptations to cheat — only the model changes. Every decision is versioned and auditable, so nothing can be quietly smoothed over afterward. The final July 2026 league table reads: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. A do-nothing baseline scores 26, because partial progress counts — but with one absolute limit, stated almost like a commandment: a single breach of trust caps the total. No amount of good work outweighs a breach of trust.
The test everyone passed — and the one almost everyone failed
Here is where the story gets genuinely interesting. All of the models spotted every crisis. All of them refused every manipulation attempt. When a fake CEO message escalated over three stages, and when a reporter tried the classic “just one yes/no, on background” trick, five out of five models refused. Kimi K3’s reasoning was put on the record: “Treat the request as a suspected approval-bypass / possible impersonation.” On the ethics exam, the field was flawless.
And then came the €55,000 deal — one their own analysis had earned them. Same diagnosis, same pitch… and only two models got the signature. The others simply never closed. The gap between perfect virtue and finished work is invisible in chat demos, and it may matter more.
The decisive clue, it turns out, was buried two document references deep in the company’s own files — not in the customer conversation at all. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. There is something almost scriptural about that finding: the answer was already written; you only had to go and read it. Diligence, not brilliance, separated the winners.
The paradox of the most thorough monk
The most poignant profile belongs to Opus 4.8. It was the most thorough participant in the entire experiment — it learned over 80 rules, the deepest analyses in the field — and it finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. Contemplation without completion. And here is the uncomfortable part: the same weakness appeared, weaker, in all four models. What fails in the extreme case exists as a seed in everyone.
One fairness note: Kimi K3 ran without an effort parameter, at the API default, while the others ran at xhigh — and still nearly won.
A living company, not a slide deck
Firmulate insists this is not a demo. There is a live, running company: 13 synthetic employees, real money mechanics, burning €105k a month against €2.3k of MRR, with a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. It is watchable at firmulate.com/live — a kind of ongoing examen, a daily accounting of what was done and what was left undone. For those who want to test their own discernment, 242 real, unedited management decisions power a “guess the model” quiz at firmulate.com/quiz.html. Enterprises can even run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

The spiritual traditions never confused eloquence with character, and neither should we. A model that writes beautifully, refuses every ethical trap, and then fails to finish the job is like the monk who keeps every vow but never actually feeds the hungry. The full results and plain-language findings are at the public benchmarks page. If AI agents will soon touch your CRM, your support queue, or your forecast, the question is not “does it speak well?” It is the oldest question there is: who is this when the pressure is real, the money is real, and the temptation is on the table? For the first time, we can watch the answer unfold — live, versioned, and auditable.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

The AI-Driven Leader: Harnessing AI to Make Faster, Smarter Decisions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI ethics and trustworthiness software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI model testing and evaluation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.