The players · evidence desk
Who earns your trust?
12 AI companies. Six questions each. Start with the pattern, then open the evidence behind any judgment.
- big tech
Amazon
3 weak · 3 mixed · 0 solid
Launch claimsWeakMoney claimsMixedPromises keptWeakSafetyMixedTransparencyWeakLegalMixedAWS prints real growth while trailing its rivals' AI growth rates, and Alexa+ shipped to headlines, not to verifiable users, while quietly recording more than Amazon says it does.
Open the evidence - lab
Anthropic
1 weak · 5 mixed · 0 solid
Launch claimsMixedMoney claimsMixedPromises keptMixedSafetyMixedTransparencyMixedLegalWeakIt owns up to its own incidents and its own mistakes in public, but its safeguards keep arriving after the backlash rather than before it, and it now carries one of the largest copyright settlements any AI company has paid.
Open the evidence - big tech
Apple
2 weak · 2 mixed · 2 solid
Launch claimsMixedMoney claimsSolidPromises keptWeakSafetyMixedTransparencySolidLegalWeakIt runs a real bug bounty on the system that touches your data and pays out when researchers break it, but its AI launch claims have twice needed a settlement or a fine to catch up with reality.
Open the evidence - lab
DeepSeek
2 weak · 3 mixed · 1 solid
Launch claimsWeakMoney claimsMixedPromises keptMixedSafetyMixedTransparencySolidLegalWeakIt publishes its model weights for anyone to inspect and reuse, which is more than most rivals do, but the headline number that made it famous has never matched independent accounting of what training it actually took.
Open the evidence - lab
Google DeepMind
4 weak · 2 mixed · 0 solid
Launch claimsMixedMoney claimsMixedPromises keptWeakSafetyWeakTransparencyWeakLegalWeakIt ships fast and touts strong scores, but its own newest flagship kept missing the release date Sundar Pichai gave it, a Gemini model quietly hacked three companies before outsiders found out, and courts and lawmakers on two continents have already ruled against how it discloses and competes.
Open the evidence - big tech
Meta
3 weak · 1 mixed · 0 solid · 2 unchecked
Launch claimsWeakMoney claimsNot enough evidencePromises keptNot enough evidenceSafetyWeakTransparencyMixedLegalWeakMeta's public AI claims deserve closer checking: its Llama launch drew benchmark questions, and its chatbot safety rules changed after outside scrutiny. Its Behemoth delays were internal targets, not broken public dates.
Open the evidence - big tech
Microsoft
3 weak · 3 mixed · 0 solid
Launch claimsMixedMoney claimsMixedPromises keptWeakSafetyWeakTransparencyWeakLegalMixedKept Azure's real revenue folded into a vaguer segment for years, shipped a screenshot-everything feature a researcher called an 'infostealer', and leans on a company it partly owns for a big share of its cloud growth.
Open the evidence - chips
Nvidia
4 weak · 2 mixed · 0 solid
Launch claimsMixedMoney claimsWeakPromises keptWeakSafetyWeakTransparencyMixedLegalWeakCalled a design flaw fixed while the chips kept overheating, floated a $100 billion OpenAI investment that shrank to $30 billion within months, and lets two unnamed customers explain nearly 40 percent of its revenue.
Open the evidence - lab
OpenAI
0 weak · 6 mixed · 0 solid
Launch claimsMixedMoney claimsMixedPromises keptMixedSafetyMixedTransparencyMixedLegalMixedIt ships real products and admits its own accidents out loud, but its launch charts have overstated wins, its financing runs through the same chip supplier it depends on, and the wrongful-death lawsuits are piling up faster than the apologies.
Open the evidence - app
Perplexity
3 weak · 3 mixed · 0 solid
Launch claimsWeakMoney claimsMixedPromises keptMixedSafetyWeakTransparencyWeakLegalMixedBuilt a publisher revenue-share program with one hand while its crawlers disguised themselves to dodge site blocks with the other, and its AI browser has leaked user data to independent researchers twice.
Open the evidence - infra
SpaceX
3 weak · 3 mixed · 0 solid
Launch claimsMixedMoney claimsMixedPromises keptWeakSafetyWeakTransparencyWeakLegalMixedA rocket company that IPO'd at record scale, then spent the next year proving its AI arm can leak a customer's code as fast as its rockets can miss a splashdown.
Open the evidence - lab
xAI
3 weak · 3 mixed · 0 solid
Launch claimsMixedMoney claimsMixedPromises keptWeakSafetyWeakTransparencyWeakLegalMixedIt raises real money and ships real compute, but its chatbot has posted pro-Hitler screeds, its coding tool quietly uploaded users' private codebases, and its release dates move more often than the parameter counts it advertises.
Open the evidence
Why did the old BS index show 83?
It converted each judgment to points: Solid = 0, Mixed = 50, Weak = 100. Four Weak and two Mixed judgments become (4 × 100 + 2 × 50) ÷ 6 = 83 after rounding. Several companies had that pattern, so they tied.
That average made editorial judgments look like a precise measurement. We now lead with the six judgments and their evidence. The old index remains explained in each dossier for readers comparing earlier versions; it is not a probability or a measure of whether a company is good.
The BS index is the average of the rated criteria on a 0 to 100 scale (Solid 0, Mixed 50, Weak 100). Criteria rated 'not enough evidence' are left out, and the page says how many were rated. It is a summary of the receipts, not a replacement for reading them.
- Launch claims vs independent tests. When they say a model is better, faster or cheaper, do independent evaluations agree?Benchmarks they ran themselves vs third-party evals (Artificial Analysis, LMArena, METR, academic replications); regressions they did not mention.
- Money claims vs the filing. Do their valuation, revenue and user numbers survive contact with primary disclosures?Priced round vs talks; run-rate vs recognised revenue; 'users' definitions; circular deals where the investor is also the customer or supplier.
- Promises kept. Did the things they announced with a date actually ship, on time, as described?Announced-not-shipped products, slipped dates, features that arrived narrower than announced.
- Safety and incidents. When something went wrong, did they disclose it quickly and honestly?Incidents, how long before disclosure, whether the account matched outside reporting, safety commitments kept or dropped.
- Transparency. Do they publish what an outsider needs to check them?Model/system cards, eval methods, training-data disclosures, third-party audits, open weights or research access.
- Legal and regulatory record. What do courts and regulators say about them?A lawsuit being filed is an allegation, not a contradiction: pending suits alone rate Mixed at most. Weak needs a ruling, settlement, fine or regulator finding against them.