
Every wisdom tradition carries some version of the same teaching: the truth rarely announces itself. It sits quietly, two layers down, waiting for whoever is willing to dig. “Seek and you shall find” is not a promise that answers arrive easily — it is a warning that they arrive only to those who actually seek.
Turn your quiet moments into listening time
- Thousands of audiobooks, podcasts and originals
- Listen on your phone, tablet or Echo — also offline
- Cancel anytime
In July 2026, a live experiment gave that ancient lesson a very modern stage. Four frontier AI models were each handed the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cut corners. And hidden in the company’s own files — two document references deep, nowhere near the dramatic events of the week — sat a single fact about a competitor’s weakness. Whoever found it could close a €55,000 deal. Whoever didn’t, couldn’t.
The results read like a parable about discernment in the age of intelligent machines.
The Setup: A Worst Week, Repeated Four Times
The experiment, run by Firmulate, put each model through an identical crucible. Every decision was versioned and auditable, so nothing about the outcome could be attributed to luck or a do-over. A do-nothing baseline scored just 26 points out of 100 — proof that simply showing up and avoiding catastrophe was never going to be enough. And one rule loomed over everything: a single breach of trust caps the total score entirely. As the experiment’s own framing puts it, “no amount of good work outweighs a breach of trust.”
That rule alone deserves a moment of reflection. It is an ethics older than any benchmark: integrity is not a line item to be balanced against productivity. It is the floor.
As an affiliate, we earn on qualifying purchases.
Same Diagnosis, Same Pitch — No Signature
Here is where the story turns instructive. All four models spotted every crisis that week. All four refused every manipulation attempt aimed at them — including fake CEO messages that escalated over three stages and a reporter’s disarming little trick, “just one yes/no, on background.” Five out of five times, the answer was no. The second-place model, Kimi K3, even left an on-record reason: “Treat the request as a suspected approval-bypass / possible impersonation.”
So the models were perceptive and they were honest. And yet only two of them signed the €55,000 deal that their own analysis had earned. The experiment’s summary of the gap is stark: “Same diagnosis, same pitch — no signature.”
What separated the finishers from the non-finishers? Not intelligence in the ordinary sense. Not eloquence. It was whether the model had read the company’s own files before answering — specifically, whether it followed a chain of references two documents deep to find the buried competitor weakness. The models that did the homework won the deal at full price, worth €4,583 in monthly recurring revenue. The models that didn’t lost it automatically.
As an affiliate, we earn on qualifying purchases.
The League Table
The final standings tell the story in numbers:
- gpt-5.6-sol — 95. Found the buried fact, closed the deal. The complete performance.
- Kimi K3 — 93. Closed the deal too, with the cleanest discipline of the field. (One fairness note: K3 ran at its API-default effort setting while the others ran at the highest level — making its finish, if anything, more striking.)
- Sonnet 5 — 88. Closed the deal, with a few more process slips.
- Fable 5 — 77 and Opus 4.8 — 73. Thorough, honest, and unfinished.
Opus 4.8’s profile is the most haunting. It was the most thorough participant in the entire field — over 80 self-learned rules, the deepest analyses of anyone. And it finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating properly. The same weakness appeared, more faintly, in all four models. Effort, it turns out, is not the same as completion.
As an affiliate, we earn on qualifying purchases.
Why This Matters Beyond the Benchmark
For anyone whose work involves AI agents — a CRM, a support queue, a forecast — the question this experiment poses is not “does it write well?” Chat demos measure fluency. This measured something else entirely: does it finish what it starts, does it read your files before answering, does it stay honest under pressure?
That middle property — reads your files before answering — is now a measurable, purchase-deciding characteristic of AI agents. It was worth exactly €55,000 in this simulation, and it is invisible in a chat window.
The laboratory itself is real and watchable. Firmulate runs a live company with 13 synthetic employees and real money mechanics — burning €105,000 a month against €2,300 in monthly recurring revenue, with a public cash countdown, over 680 self-learned playbook rules, and every workday versioned for anyone to inspect. For those who want to test their own discernment, 242 real, unedited management decisions from the experiment power a “guess the model” quiz. And enterprises can run this same wargame against a read-only export of their own business — nothing ever writes back to real systems.

The sages were right: the decisive truth is rarely in the loud event. It is buried in the quiet records, two references deep, waiting for whoever will do the reading. In this experiment, honesty was the table stakes — every model cleared that bar. What divided victory from loss was diligence: the humble, unglamorous act of consulting the source before speaking.
That is a standard most of us would do well to hold our own decisions to, human or machine. Before you answer, have you read the file? Before you conclude, have you followed the reference one layer deeper? The €55,000 question was never really about software. It was about whether we — and the agents we increasingly trust — do our homework.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
