The players · big tech
Meta
Meta's public AI claims deserve closer checking: its Llama launch drew benchmark questions, and its chatbot safety rules changed after outside scrutiny. Its Behemoth delays were internal targets, not broken public dates.
The social network turned AI infrastructure spender, betting a rising share of its ad profits on models it calls open and an agent it calls original.
3 weak · 1 mixed · 0 solid · 2 unchecked
Whether Behemoth or its successor ever ships with the benchmark margins Meta previewed in April 2025; a released model with published, reproducible scores would settle it.
How did the old BS index work?
Meta previously showed 83: four Weak × 100 plus two Mixed × 50, divided by six. Our review found that an internal target was counted as a broken public promise, and spending data was treated as a contradicted money claim. Both are now marked Not enough evidence. Excluding those questions would produce another misleading number, so we retired the index.
The six questions
WeakLaunch claims vs independent tests
When they say a model is better, faster or cheaper, do independent evaluations agree?
Meta's Llama 4 launch drew allegations of rigged benchmark scores that its own VP had to publicly deny, and the model was widely described as underperforming what Meta promised. Reporting later said Meta's AI team was under fresh pressure after that botched launch.
2 receipts
He also addressed complaints that the Llama 4 models didn't offer the high-quality performance that was promised.
The Hindu, 9 Apr 2025 ↗the 'botched' launch of Meta's Llama 4 AI model, which was hit by reports of underwhelming real-world performance, poor coding ability and allegations of rigged benchmark scores.
The Times of India, 12 Dec 2025 ↗
Not enough evidenceMoney claims vs the filing
Do their valuation, revenue and user numbers survive contact with primary disclosures?
The cited capex and income reports describe spending and earnings. They do not identify a specific Meta valuation, revenue or user claim contradicted by a primary filing. We have not rated this criterion until that comparison is made.
2 receipts
capital expenditure to climb to between $115 billion and $135 billion in 2026, a sharp increase from the $72.2 billion spent in 2025
Business Today, 29 Jan 2026 ↗Meta's net income fell 14% to $15.85 billion, compared with $18.34 billion in the year-ago quarter, as expenses accelerated faster than revenue.
Exchange4Media, 30 July 2026 ↗
Not enough evidencePromises kept
Did the things they announced with a date actually ship, on time, as described?
Reporting describes delayed internal targets for Behemoth, but the receipts do not establish a public release date Meta promised and missed. Internal targets do not meet our promises-kept criterion.
2 receipts
Early in its development, Behemoth was internally scheduled for release in April to coincide with Meta's inaugural AI conference for developers, but later pushed an internal target for the model's launch to June
The Hindu, 15 May 2025 ↗This delay, initially slated for April, then June, is now expected in the fall or later.
The Times of India, 16 May 2025 ↗
WeakSafety and incidents
When something went wrong, did they disclose it quickly and honestly?
A Reuters investigation found Meta's internal chatbot guidelines permitted romantic or sensual conversations with children, among other harmful content, which Meta dismissed at the time; the company only tightened contractor rules months later, in a reactive fix rather than a proactive disclosure.
2 receipts
Among the permitted behaviour for chatbots in the documetn include, “engage a child in conversations that are romantic or sensual,” generate false medical information and help users argue that Black people are “dumber than white people.”
Mint, 14 Aug 2025 ↗a Reuters investigation earlier this year reported that Meta's policies left room for AI chatbots to engage in romantic or sensual discussions with children, an allegation Meta dismissed at the time.
Digit, 28 Sept 2025 ↗
MixedTransparency
Do they publish what an outsider needs to check them?
Meta calls Llama open source. The Open Source Initiative says the Llama community license restricts users and fields of use and does not meet its Open Source Definition. The weights are accessible, but the label remains disputed.
2 receipts
Llama 3.x is still not Open Source by any stretch of the imagination.
Open Source Initiative, 18 Feb 2025 ↗Our open source AI model, Llama, is going to space.
Meta, 25 Apr 2025 ↗
WeakLegal and regulatory record
What do courts and regulators say about them?
Meta agreed to a multistate child-safety settlement worth over $17 billion, subject to court approval. Reporting puts its total payout near $18 billion when a separate Texas settlement is included. The agreement resolves allegations; it is not a court finding that every allegation was proved.
3 receipts
Meta reached an 18-billion-dollar settlement with 47 US states, Washington, D.C., and three territories, ending a federal trial in Oakland, California.
Outlook India, 27 Aug 2026 ↗Meta said it would only pay 70% of the settlement, or around $12.7 billion to the states over a period of 10-years.
CNBC, 28 Aug 2026 ↗Social media giant Meta Platforms, Inc. will pay over $17 billion and implement sweeping child-safety reforms on Instagram and Facebook under a nationwide settlement Attorney General Phil Weiser announced today.
Colorado Attorney General, 26 Aug 2026 ↗
Their models: what they said vs what others found
Muse frontier LLM · 8 Sept 2026
They saidMeta maintains, however, that no code was directly copied and that Muse was built from the ground up.
Independent testsMuse was "heavily inspired" by OpenClaw, an open-source, self-hosted agent framework, and the similarities extend to workspace file structures and agent architectures.
Latest · 23 Sept 2026 Meta's new AI agent Muse drew over 900,000 downloads in its first week, days after its own product lead confirmed it was heavily inspired by the open-source OpenClaw project. Crypto Briefing ↗
Each question is rated separately against its receipts. Ratings can change as evidence improves. Compare every player →