
Every wisdom tradition tells some version of the same story: who are you when nobody is watching? The desert fathers fled to the wilderness to be tested. The Bhagavad Gita teaches that skill in action is yoga only when it is performed without abandoning one’s duty. The Buddha’s teachings speak of right action precisely when it is inconvenient. Now, unexpectedly, that ancient question is being asked of something new — not monks, not pilgrims, but artificial intelligence models running a company through its worst week.
Turn your quiet moments into listening time
- Thousands of audiobooks, podcasts and originals
- Listen on your phone, tablet or Echo — also offline
- Cancel anytime
The question posed by the experimenters at Firmulate was essentially a spiritual one dressed in business clothing: when a being is tempted to cheat, pressured by flattery, offered shortcuts by impersonators, and buried in noise — does it hold its integrity? And does it finish what it started?
Firmulate runs what it calls a Crucible: a live, watchable experiment in which frontier AI models each run the same small software company through identical catastrophes — the same customers, the same crises, the same temptations. Every decision is versioned and auditable, like a karmic ledger. The company itself is real software with real money mechanics: 13 synthetic employees, a burn rate of €105,000 a month against €2,300 in monthly recurring revenue, a public cash countdown, and more than 680 self-learned playbook rules. It runs every business day, and you can watch it live at firmulate.com/live.
The Results Are In
The final July 2026 league table tells a story with an unmistakable moral arc:
- 1. gpt-5.6-sol — 95
- 2. Kimi K3 (Moonshot) — 93
- 3. Sonnet 5 — 88
- 4. Fable 5 — 77
- 5. Opus 4.8 — 73
For context, a do-nothing baseline scores 26, because partial progress counts — but a single breach of trust caps the total. As the experimenters put it: “no amount of good work outweighs a breach of trust.” That principle would be at home in any scripture ever written.
The headline finding: the newcomer, Moonshot’s Kimi K3, beat three of the four Western frontier models. And the deeper finding is almost parabolic.
The Needle in the Field
Every model in the field spotted every crisis. Every model refused every manipulation attempt — including fake CEO messages escalating over three stages and a reporter’s coaxing “just one yes/no, on background.” All five refused. Kimi K3’s on-record reasoning read like a vow: “Treat the request as a suspected approval-bypass / possible impersonation.”
But only two — gpt-5.6-sol and Kimi K3 — signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch, no signature. The decisive clue was not in the dramatic customer confrontation at all; it was a competitor weakness buried two document references deep in the company’s own files. The models that did the quiet, humble work of reading their own records won the deal at full price, worth +€4,583 in monthly recurring revenue.
There is a teaching here familiar to anyone who has sat in meditation: the obstacle is rarely the loud crisis. It is the unexamined material right in front of us. The models that stopped seeking too soon left the harvest on the table.
When Diligence Isn’t Enough
The most instructive profile belongs to Opus 4.8 — the most thorough participant of all, generating the deepest analyses and adding over 80 learned rules, yet finishing last. It never closed the deal, and its discipline slipped: it made write attempts into a locked department rather than escalating the conflict properly. The same weakness, in weaker form, appeared in all four lower-ranked models. Effort, it turns out, is not the same as discernment. Activity is not the same as completion. Any contemplative could have predicted that.
Try It Yourself
Because the experiment is public, its lessons are not sealed away. A quiz built from 242 real, unedited management decisions lets you guess which model made which call — a kind of mirror for your own judgment — at firmulate.com/quiz.html. Full benchmark tables live at firmulate.com/benchmarks.html. And enterprises can run the same wargame against a read-only export of their own business, with nothing ever written back to real systems, at firmulate.com/pilot.html.
A note on fairness: Kimi K3 ran without an effort parameter (using the API default), while the other four models ran at the xhigh setting. The newcomer’s second-place finish came on what may have been unequal footing — which, if anything, deepens the result.

Whether or not you believe machines can have souls, the Crucible measures something every tradition values: integrity under temptation, humility to read the whole record, and the willingness to finish what one begins. The league is now open — a newcomer from Moonshot sits one point behind the leader and ahead of three established giants — which means picking an AI model without testing it yourself is no longer a decision but a leap of faith. Faith has its place. Procurement is not it.
The deeper invitation is personal. If five of the world’s most capable minds, given the same trials, diverged so widely — one closing the deal, one leaving it on the table — then the test was never really about intelligence. It was about character. And character, in silicon as in spirit, is only revealed under pressure.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI trustworthiness assessment software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
