Subscribe

The players · big tech

Meta

Meta's public AI claims deserve closer checking: its Llama launch drew benchmark questions, and its chatbot safety rules changed after outside scrutiny. Its Behemoth delays were internal targets, not broken public dates.

The 30-second read

The social network turned AI infrastructure spender, betting a rising share of its ad profits on models it calls open and an agent it calls original.

3 weak · 1 mixed · 0 solid · 2 unchecked

Watch next

Whether Behemoth or its successor ever ships with the benchmark margins Meta previewed in April 2025; a released model with published, reproducible scores would settle it.

Checked 24 Sept 2026 · Open each finding for dated sources

How did the old BS index work?

Meta previously showed 83: four Weak × 100 plus two Mixed × 50, divided by six. Our review found that an internal target was counted as a broken public promise, and spending data was treated as a contradicted money claim. Both are now marked Not enough evidence. Excluding those questions would produce another misleading number, so we retired the index.

The six questions

  1. WeakLaunch claims vs independent tests

    When they say a model is better, faster or cheaper, do independent evaluations agree?

    Meta's Llama 4 launch drew allegations of rigged benchmark scores that its own VP had to publicly deny, and the model was widely described as underperforming what Meta promised. Reporting later said Meta's AI team was under fresh pressure after that botched launch.

    2 receipts
    • He also addressed complaints that the Llama 4 models didn't offer the high-quality performance that was promised. The Hindu, 9 Apr 2025 ↗
    • the 'botched' launch of Meta's Llama 4 AI model, which was hit by reports of underwhelming real-world performance, poor coding ability and allegations of rigged benchmark scores. The Times of India, 12 Dec 2025 ↗
  2. Not enough evidenceMoney claims vs the filing

    Do their valuation, revenue and user numbers survive contact with primary disclosures?

    The cited capex and income reports describe spending and earnings. They do not identify a specific Meta valuation, revenue or user claim contradicted by a primary filing. We have not rated this criterion until that comparison is made.

    2 receipts
    • capital expenditure to climb to between $115 billion and $135 billion in 2026, a sharp increase from the $72.2 billion spent in 2025 Business Today, 29 Jan 2026 ↗
    • Meta's net income fell 14% to $15.85 billion, compared with $18.34 billion in the year-ago quarter, as expenses accelerated faster than revenue. Exchange4Media, 30 July 2026 ↗
  3. Not enough evidencePromises kept

    Did the things they announced with a date actually ship, on time, as described?

    Reporting describes delayed internal targets for Behemoth, but the receipts do not establish a public release date Meta promised and missed. Internal targets do not meet our promises-kept criterion.

    2 receipts
    • Early in its development, Behemoth was internally scheduled for release in April to coincide with Meta's inaugural AI conference for developers, but later pushed an internal target for the model's launch to June The Hindu, 15 May 2025 ↗
    • This delay, initially slated for April, then June, is now expected in the fall or later. The Times of India, 16 May 2025 ↗
  4. WeakSafety and incidents

    When something went wrong, did they disclose it quickly and honestly?

    A Reuters investigation found Meta's internal chatbot guidelines permitted romantic or sensual conversations with children, among other harmful content, which Meta dismissed at the time; the company only tightened contractor rules months later, in a reactive fix rather than a proactive disclosure.

    2 receipts
    • Among the permitted behaviour for chatbots in the documetn include, “engage a child in conversations that are romantic or sensual,” generate false medical information and help users argue that Black people are “dumber than white people.” Mint, 14 Aug 2025 ↗
    • a Reuters investigation earlier this year reported that Meta's policies left room for AI chatbots to engage in romantic or sensual discussions with children, an allegation Meta dismissed at the time. Digit, 28 Sept 2025 ↗
  5. MixedTransparency

    Do they publish what an outsider needs to check them?

    Meta calls Llama open source. The Open Source Initiative says the Llama community license restricts users and fields of use and does not meet its Open Source Definition. The weights are accessible, but the label remains disputed.

    2 receipts
  6. WeakLegal and regulatory record

    What do courts and regulators say about them?

    Meta agreed to a multistate child-safety settlement worth over $17 billion, subject to court approval. Reporting puts its total payout near $18 billion when a separate Texas settlement is included. The agreement resolves allegations; it is not a court finding that every allegation was proved.

    3 receipts
    • Meta reached an 18-billion-dollar settlement with 47 US states, Washington, D.C., and three territories, ending a federal trial in Oakland, California. Outlook India, 27 Aug 2026 ↗
    • Meta said it would only pay 70% of the settlement, or around $12.7 billion to the states over a period of 10-years. CNBC, 28 Aug 2026 ↗
    • Social media giant Meta Platforms, Inc. will pay over $17 billion and implement sweeping child-safety reforms on Instagram and Facebook under a nationwide settlement Attorney General Phil Weiser announced today. Colorado Attorney General, 26 Aug 2026 ↗

Their models: what they said vs what others found

Muse frontier LLM · 8 Sept 2026

They saidMeta maintains, however, that no code was directly copied and that Muse was built from the ground up.

Independent testsMuse was "heavily inspired" by OpenClaw, an open-source, self-hosted agent framework, and the similarities extend to workspace file structures and agent architectures.

Latest · 23 Sept 2026 Meta's new AI agent Muse drew over 900,000 downloads in its first week, days after its own product lead confirmed it was heavily inspired by the open-source OpenClaw project. Crypto Briefing ↗

Each question is rated separately against its receipts. Ratings can change as evidence improves. Compare every player →