firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Imagine training for a marathon where your coach not only guides your steps but also decides when to push harder or pull back, all while staying honest and focused on your goal. In the world of AI, a similar challenge unfolds: can these models actually run a business — stay disciplined, reliable, and honest — under pressure? The answer isn’t just about how well they chat but whether they can deliver real results when it counts. This question is now being tested in a groundbreaking experiment that reveals startling insights for everyone, even those focused on fitness and self-improvement.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get workout gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Business of Testing AI — Like a Fitness Regimen

Just as athletes undergo rigorous training to prepare for their biggest races, AI models are put through demanding tests to evaluate their ability to handle real-world challenges. Recently, a live experiment called the Crucible League brought together five leading AI models, including a newcomer called Kimi K3 from Moonshot, to run a small software company through its worst week. This wasn’t a simple chat test — it was a full-blown simulation involving crises, customer demands, and ethical dilemmas.

All models faced identical scenarios, with the same customers, the same crises, and the same temptations to cheat or cut corners. Every decision was recorded and auditable, ensuring fairness and transparency. The goal? To see which AI could truly manage a company, make honest decisions, and close a critical deal worth €55,000 — about $58,000 USD — and an increase of €4,583 in monthly recurring revenue.

The Surprising Results

  • The top performer, gpt-5.6-sol, scored 95 out of 100, found a buried piece of critical information in the company’s files, and closed the deal. It demonstrated complete awareness and discipline.
  • Kimi K3, a relative newcomer, scored just slightly behind at 93, but what truly stood out was its discipline. It refused all manipulative tactics, read deeply into the company’s documents, and closed the deal—earning the full €55,000 deal and boosting monthly revenue by €4,583.
  • Other models, such as Sonnet 5, scored 88 and 77, respectively. They also closed the deal but showed more slips and process lapses, such as leaving the close on the table or escalating issues into locked departments.
Amazon

fitness tracking smartwatch

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Lessons from the Experiment — Discipline and Integrity Are Key

This experiment reveals a vital truth: AI’s ability to stay disciplined under pressure and to read and interpret the company’s internal files deeply was crucial in winning the deal. Interestingly, the weakness of competing models was not in spotting crises but in their failure to dig into the company’s documents to find the buried facts that clinched the deal.

Also notable is what didn’t happen: all models, including those that scored lower, refused attempts at social engineering and manipulation, such as fake CEO messages and reporter tricks. Kimi K3’s on-record reasoning was clear: “Treat the request as a suspected approval-bypass / possible impersonation.” This shows that ethical behavior under pressure is achievable and measurable.

The Real Company — Live and Growing

Behind the experiment is a real, functioning company with 13 synthetic employees managing real money mechanics, burning €105k monthly against a mere €2.3k in MRR. The system is public, transparent, and live, providing ongoing data and decision-making challenges every workday. Watch the company in action at firmulate.com/live.

This setup allows enterprises and individuals to run similar wargames against their own business models, testing whether their AI helpers can stay honest, diligent, and effective before making any real-world commitments.

Amazon

high-precision running GPS watch

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Takeaway — Choosing the Right AI Model Is a Critical Business Decision

For fitness enthusiasts and self-improvement advocates, the lesson extends beyond business: discipline, focus, and integrity are essential whether you’re training for a marathon or trying to build better habits. Just as an AI that reads deeply and refuses shortcuts performs better in a business setting, a person who stays disciplined and honest under pressure will reach their goals more reliably.

As AI continues to integrate into your work and personal life, understanding which models can finish what they start — read the files, resist temptations, and deliver real results — becomes crucial. The league table from this experiment places gpt-5.6-sol at the top, with Kimi K3 closely behind, demonstrating that even newcomers can excel with the right discipline.

Amazon

performance monitoring fitness band

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Fairness Note

It’s important to mention that Kimi K3 ran without an effort parameter (API default), while the other models ran at xhigh. This ensures a fair comparison, emphasizing discipline and decision quality over effort settings.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

AI models that stay disciplined, honest, and deeply read your internal data outperform those that slip or cut corners—much like a focused athlete. Choosing the right AI is a vital business decision, just as choosing your training approach determines your success in fitness.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


Amazon

advanced running shoes for marathon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How AI’s Deep File Reading Wins Big Deals — and What It Means for Business Decisions

Discover how AI models that read deeply into company files outperform others by uncovering hidden facts, winning deals, and making smarter business decisions.