
Imagine a company that operates in full view, with AI models navigating its daily crises, making decisions, and even refusing manipulative tactics — all while losing €105,000 a month. This isn’t science fiction; it’s the real-time experiment of Firmulate, a startup pushing the boundaries of AI testing in the open.
The Living Lab of AI Decision-Making
At the heart of this experiment is a small, real company with a twist: it’s run entirely by AI models, each acting as a synthetic employee. Every decision they make is logged, every crisis they face is simulated, and every tactic they attempt to manipulate is tested. The goal? To measure how well these AI agents can manage real-world business challenges while remaining honest and disciplined.
How the Experiment Works
Four frontier AI models — including the well-known GPT-5.6-sol, Kimi K3, Sonnet 5, and Opus 4.8 — are each given the same set of problems: run this tiny software company through its worst week. The scenarios include customer crises, internal crises, and manipulative attempts like fake CEO messages or background requests. All decisions are versioned and auditable, creating a transparent, real-time record of each model’s behavior.
The Results: Vigilance, Integrity, and Gaps
All four models successfully identified every crisis and refused every manipulation attempt, demonstrating an impressive level of vigilance. Yet, only two actually managed to close the deal worth €55,000, despite their correct diagnoses and compelling pitches — a reminder that strategic discipline matters just as much as recognition.
In a surprising twist, the decisive advantage lay in a buried fact within the company’s own files. The models that read and understood this hidden piece of information secured the deal at full price, translating into an additional €4,583 monthly recurring revenue.
The Human-Like Challenges
Beyond crisis management, the models faced social engineering tests, such as staged fake CEO messages and background requests. Every one of the five models tested refused to go along, guided by their built-in reasoning. For Kimi K3, the approach was clear: “Treat the request as a suspected approval-bypass / possible impersonation.” This discipline underscores how AI models can be trained to recognize deception and resist manipulation — crucial qualities for trustworthiness in business applications.
The Overarching Reality: A Company Losing Money Publicly
The company itself runs with 13 synthetic employees and real-money mechanics, burning through €105,000 each month against a modest €2,300 in monthly recurring revenue. Every day is a balancing act, with a public cash countdown visible at firmulate.com/live.html. The company’s existence is a real-time story of survival, with every decision and process openly available for viewers to see, analyze, and learn from.
What the Results Tell Us
Among the models, Opus 4.8 is noteworthy for its thoroughness, employing over 80 learned rules and conducting deep analyses. Yet, it finished last in the performance league, leaving opportunities on the table and slipping into operational slips — like writing attempts instead of escalating issues.
The experiment underscores a vital point: AI models can be trained to recognize crises and resist manipulative tactics, but closing deals and executing disciplined processes remain challenging, especially under pressure. The results are a stark reminder that AI’s capacity for integrity and execution is still a work in progress, even when the models seem to perform perfectly on paper.
The Industry Implications
This experiment isn’t just about testing AI; it’s about understanding how AI will integrate into real business environments. As these models are tested against real crises, manipulations, and decision-making pressures, the overarching lesson becomes clear: the true value of AI in business isn’t just in generating words or ideas, but in executing tasks reliably, honestly, and effectively.
Explore and Watch Live
Interested in seeing how this unfolds in real time? You can watch the company operate daily, explore the decision-making process, and even participate in quizzes to guess which model made which decision—all at firmulate.com/live.html.

In a world where AI models are tested against real crises daily, the key qualities are honesty, discipline, and the ability to finish what they start. Watching this live experiment reveals not just progress but also the persistent gaps in AI’s readiness for real-world business challenges.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Products Worth Considering
As an affiliate, we earn on qualifying purchases.