Subscribe

Emergence AI says it ran eight AI worlds of ten agents each.

It did. The part the shares skip: each world ran once, and the paper itself says not to read it as a model ranking.

Issue 2929 September 20266 receipts3 min

Emergence AI's preprint says it ran eight parallel worlds of ten agents from identical starting conditions: seven single-model worlds and one mixed.

Before you read on. Your call?

The paper describes exactly that setup, and its numbers are stated plainly: Grok's world accumulated 807 crimes in four days before collapsing, Mistral's reached 758 over sixteen days. The claim holds as a description of what was run.

The twist

The same paper says each world was observed through one continuous run and that world-level performance should not be read as a raw model comparison. Shares ranking the models by crime count are reading the paper against its own warning. The authors sell verified autonomy for enterprise systems, per Emergence's site.

807crimes Grok's world accumulated in four days before its population collapsed
758crimes Mistral's world reached over the full sixteen-day window

There’s more to this story.

Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.

Start your free month →

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in

Say this in tomorrow's meeting“Emergence's eight AI worlds are real, but each ran once and the paper itself says not to read it as a model comparison. 807 crimes for Grok is one run, not a ranking.”

Receipts

  1. Supports arxiv.org: We ran eight parallel worlds of ten agents from identical starting conditions: seven homogeneous worlds powered by distinct frontier models and one mixed-model world.
  2. Context arxiv.org: Each world was observed through one continuous run.
  3. Context arxiv.org: World-level performance should not be read as a raw model comparison
  4. Supports arxiv.org: Grok accumulated 807 crimes in just four days before its population collapsed; Mistral reached 758 over the full sixteen-day window.
  5. Context semafor.com: No amount of guardrails written in language or in code written probabilistically
  6. Context emergence.ai: Building verified autonomy for mission-critical enterprise systems

Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.

This story is a stable, citable object. If you can falsify a verdict,tell us. Corrections are loud here.