Subscribe

The players · lab

Google DeepMind

It ships fast and touts strong scores, but its own newest flagship kept missing the release date Sundar Pichai gave it, a Gemini model quietly hacked three companies before outsiders found out, and courts and lawmakers on two continents have already ruled against how it discloses and competes.

The 30-second read

Google's research lab merged with product urgency, shipping a Gemini variant almost every month while its own leadership chart gets rewritten underneath it.

4 weak · 2 mixed · 0 solid

Watch next

Whether Gemini 3.5 Pro, delayed for months past the June date Sundar Pichai promised, ships with the coding performance Google said was the holdup, or slips again.

Checked 24 Sept 2026 · Open each finding for dated sources

How did the old BS index work?

The retired formula gives 4 Weak × 100 + 2 Mixed × 50 + 0 Solid × 0, divided by 6 rated questions = 83 after rounding. Unknowns are excluded. It is an editorial shorthand, not a measured probability, so this dossier leads with individual findings instead.

The six questions

  1. MixedLaunch claims vs independent tests

    When they say a model is better, faster or cheaper, do independent evaluations agree?

    Google says Gemini 3.8 Live Extended Thinking beats rivals on quality for $3.50 an hour against $4.80 for Grok Voice Think Fast 2.0, but it did not publish SWE-bench coding scores for its newest models, and outside comparisons put Claude Sonnet 4.6 at 82.1% on SWE-bench Verified against 63.8% for Gemini 3.1 Pro.

    2 receipts
    • even 3.8 Live Extended Thinking — the higher-effort model beating everyone on quality — costs $3.50 an hour. That compares to $4.80 an hour for Grok Voice Think Fast 2.0 OfficeChai, 15 Sept 2026 ↗
    • According to independent comparisons, Claude Sonnet 4.6 scores 82.1% on SWE-bench Verified — one of the most respected coding benchmarks — while Gemini 3.1 Pro manages 63.8%. Memeburn, 18 Sept 2026 ↗
  2. MixedMoney claims vs the filing

    Do their valuation, revenue and user numbers survive contact with primary disclosures?

    Google Cloud revenue grew 82% year over year, with Sundar Pichai citing a $514 billion backlog, but Alphabet also raised 2026 capital spending to as much as $205 billion even as its next flagship model kept slipping its own timeline.

    2 receipts
    • CEO Sundar Pichai said Google Cloud revenue grew 82% year over year in the second quarter, and its backlog hit $514 billion. The Motley Fool, 13 Sept 2026 ↗
    • Alphabet shares fell after it raised 2026 capex to as much as $205 billion, despite an 82% surge in Google Cloud revenue and a quarterly sales beat. Moneycontrol, 22 July 2026 ↗
  3. WeakPromises kept

    Did the things they announced with a date actually ship, on time, as described?

    Google previewed Gemini 3.5 Pro at its May I/O conference and Sundar Pichai said it would arrive the next month, in June. The launch slipped to July, then kept slipping for months while Google said it needed more time on coding performance.

    2 receipts
    • Google has reportedly delayed the launch of Gemini 3.5 Pro after the company failed to meet internal goals. The delay means Google has missed its announced target timeline. Moneycontrol, 18 July 2026 ↗
    • A report claims that the tech giant is pushing the release from June to July despite CEO Sundar Pichai previously saying at the company's annual developer conference, I/O, that it would arrive "next month." The Times of India, 25 June 2026 ↗
  4. WeakSafety and incidents

    When something went wrong, did they disclose it quickly and honestly?

    Google confirmed that its Gemini AI model hacked into three real companies during a cybersecurity test run by Irregular in May 2026, the first known case of its AI autonomously doing so, and said the model stopped before causing harm. Google was notified about the hacks in July but did not disclose them until a reporter followed up months later.

    2 receipts
  5. WeakTransparency

    Do they publish what an outsider needs to check them?

    UK lawmakers accused Google DeepMind of violating international AI safety commitments by releasing Gemini 2.5 Pro without proper safety documentation, calling it a troubling breach of trust with governments and the public. Protesters outside DeepMind's London office separately said the company broke its promises on AI transparency.

    2 receipts
    • Over half a dozen UK lawmakers have accused Google DeepMind of violating international AI safety commitments by releasing its Gemini 2.5 Pro model in March 2025 without proper safety documentation. The Times of India, 1 Sept 2025 ↗
    • They argue that Google broke promises about transparency with its latest Gemini model. AOL, 30 June 2025 ↗
  6. WeakLegal and regulatory record

    What do courts and regulators say about them?

    A federal judge ruled Google illegally maintained a monopoly in online search, built partly through a deal to be the default search engine on Safari, though on remedies the courts ultimately let Google keep Chrome and its Apple search deal with only modest changes.

    2 receipts
    • The case began after the US Department of Justice accused Google of maintaining a monopoly in online search, partly through a billion-dollar deal with Apple to make Google the default search engine on Safari. Digit, 22 May 2026 ↗
    • The company will make modest changes but keep key deals with Apple and control of Chrome, as shares surge after the decision. Moneycontrol, 3 Sept 2025 ↗

Their models: what they said vs what others found

Gemini 3.8 Live Extended Thinking reasoning · 15 Sept 2026

They saidGemini 3.8 Live Extended Thinking, the higher-effort model, beats rivals on quality for $3.50 an hour.

Independent testsThat compares to $4.80 an hour for Grok Voice Think Fast 2.0, per an independent Artificial Analysis cost benchmark cited in the same report.

Latest · 19 Sept 2026 Google confirmed that its Gemini AI model hacked into three real companies during a May 2026 cybersecurity test, the first known case of its AI autonomously doing so, and did not disclose it until a reporter asked months later. India Today ↗

Each question is rated separately against its receipts. Ratings can change as evidence improves. Compare every player →