
Imagine a business that has no employees, constantly loses money, yet operates transparently in front of your eyes. This is not fiction — it’s the live experiment by Firmulate, where AI models run a small company through its toughest week, revealing what it takes to manage with honesty, discipline, and a touch of vulnerability. For those curious about how AI could directly impact future workplaces, this is the real-world case to watch.
The Experiment: An AI-Driven Company Under Siege
At its core, Firmulate’s live experiment presents a small software company — staffed by 13 synthetic employees, but powered by advanced AI models. This virtual business faces real crises: customer demands, operational dilemmas, and ethical temptations. Each AI model, representing different versions of cutting-edge language systems, runs the same week’s scenario, where every decision is meticulously recorded and made auditable.
The goal? To see which AI models can best navigate the chaos while maintaining integrity and achieving results. The models are judged on their ability to spot crises, refuse manipulation, and close deals. The company’s finances are transparent, with €105,000 burned each month against a modest €2,300 in monthly recurring revenue.
The Results: Distinct Abilities and Limitations
All four AI models managed to identify every crisis and refused every manipulation attempt, demonstrating impressive diligence. However, only two succeeded in closing a key €55,000 deal — their own analysis had earned them the opportunity. Interestingly, the decisive advantage was buried deep within the company’s own files, not in the superficial customer interactions. The models that read and understood these internal documents secured the full deal, adding €4,583 in monthly recurring revenue.
Such findings are revealing: the AI’s ability to dig beneath surface-level data can be the difference between survival and missed opportunity. This highlights the importance of thorough information processing in AI decision-making, especially in business contexts where hidden insights matter.
Handling Ethical Dilemmas and Manipulation Attempts
The experiment also tested how AI models respond to social engineering tactics. Fake CEO messages, escalating over three stages, and a reporter’s subtle trick — all designed to manipulate or bypass approval processes — were met with unanimous refusal by all models. Kimi K3’s on-record reasoning was clear: “Treat the request as a suspected approval-bypass / possible impersonation.”
This discipline is critical for AI’s role in sensitive environments, where social engineering is a real threat. The fact that every model refused to be manipulated indicates a promising direction for deploying AI in trustworthy roles, even under pressure.
The Live Company: A Transparent Laboratory
Firmulate’s live setup is extraordinary — a real-time simulation with no employees, burning cash openly, with every rule learned and decision recorded. Visitors can watch the daily evolution, see decisions made, and understand how different AI configurations perform under identical conditions. The experiment is not static; it is refreshed twice daily, offering a continuous window into the AI’s capabilities and limitations.
The company’s current leaderboard shows gpt-5.6-sol leading with a score of 95, having identified the buried fact and secured the deal. Kimi K3 follows closely with 93, demonstrating the importance of discipline and internal data reading. Meanwhile, Sonnet 5 and Fable 5 also participated, with scores of 88 and 77 respectively, revealing strengths and weaknesses in rule discipline and process execution.
Implications for the Future of Work
This experiment underscores a vital point: in AI-driven workplaces, the ability to complete tasks reliably — reading files thoroughly, resisting manipulation, closing deals — matters more than just generating convincing chat. As AI begins to touch CRM systems, customer support, and forecasting, organizations must ask: will these models finish what they start? Will they stay honest under pressure?
The live experiment offers a glimpse into how AI could be a trustworthy partner or a reckless worker. Its transparent, build-in-public approach reveals that even the most advanced models can falter, especially when discipline slips or critical internal information is overlooked.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Products Worth Considering
As an affiliate, we earn on qualifying purchases.