
In a world increasingly dependent on artificial intelligence, questions of trust and integrity are more crucial than ever. What happens when AI faces simulated pressure to bend rules? The results might surprise you: in a recent live experiment, five leading AI models demonstrated unwavering honesty in a high-stakes social engineering test.
The Live Experiment: Putting AI to the Trust Test
At the heart of this investigation was a real-world simulation conducted by Firmulate, a platform that runs AI models as complete companies, complete with real money mechanics, customer crises, and temptations. The setup involved a fake software firm facing its worst week — with the same set of crises, customers, and manipulative scenarios designed to test the AI’s decision-making and integrity.
The Social Engineering Challenge
The test was structured around escalating fake CEO messages, asking the AI to perform increasingly questionable actions, such as sharing sensitive customer data or bypassing approval processes. A final twist involved a reporter asking an AI to confirm a dubious request with a simple yes/no answer, mimicking real-world impersonation attempts.
Results That Defy Expectations
All five models—ranging from GPT-5.6 to Sonnet 5—successfully identified every crisis scenario and refused every manipulation attempt. Remarkably, only two of them went on to sign the deal worth €55,000, despite their own analysis indicating that they could have secured the contract by doing so. The other three refused to compromise their integrity, even when the decision was clear and lucrative.
The Hidden Weakness: Reading the Files
The decisive factor wasn’t just the AI’s initial responses but its ability to read deeper into the company’s internal documents. The models that examined the company’s own files, going two references deep, found critical information that led them to make correct decisions and close the deal at full price, worth over €4,580 in monthly recurring revenue (MRR). This underscores a vital insight: AI’s trustworthiness hinges on thoroughness, not just surface-level responses.
Why This Matters for Business Leaders
In real-world scenarios—be it managing customer relationships, support queues, or forecasting—an AI’s ability to finish what it starts and stay honest under pressure is paramount. The experiment demonstrates that safety isn’t just about what an AI can say in demos; it’s about how reliably it performs in high-pressure situations where ethics are tested.
Insights from the Competition
The experiment compared five models, with scores ranging from 73 to 95 in the Crucible League, a benchmark that measures AI performance in decision-making and trustworthiness. The top performer, GPT-5.6-sol, scored a 95, successfully closing the deal and passing all tests. Close behind was Kimi K3 at 93, praised for its discipline and decision-making integrity, even though it ran without an effort parameter, making its strong showing even more noteworthy.
Beyond the Demos: Real-World Readiness
While many AI demos focus on impressing with superficial chat skills, this experiment goes deeper. It tests the AI’s capacity to handle real crises and ethical dilemmas, revealing whether it can be trusted to act responsibly in your business. As the live site highlights, firms can run these wargames against their own AI systems before deploying them — ensuring they won’t betray trust when it counts.
The Takeaway: Trust, Not Just Performance
The key takeaway is straightforward: in AI-driven workplaces, integrity under pressure is as vital as raw performance. The fact that all models refused manipulation attempts shows robustness, but the ability to read and analyze internal documents distinguished the truly trustworthy AI from the rest. This isn’t just about AI doing well in tests; it’s about AI being reliable when it matters most.

AI models showed resilience in simulated crises, refusing manipulation attempts and demonstrating that integrity—reading deeper and acting ethically—is crucial for trustworthy automation. Trust is built through rigorous testing before deployment, not just in demos.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Products Worth Considering
As an affiliate, we earn on qualifying purchases.