Learn · a trick with a name
Cherry-Picked Slice
the flattering subset, presented as the whole.
The best segment of a result, shown alone, looks like the whole result.
How to spot it
- Results for one region, cohort or task only
- Averages that quietly exclude the worst cases
- Charts that start or stop at a convenient date
The one question to ask“What does the full set look like?”
At work
When a vendor shows a case study, ask for the median customer, not the best one.
Caught in the wild
Every time we've caught it so far. Try calling each one before you open it.
- TRUE, BUTA new NBER survey of 6,000 executives says 9 in 10 saw zero AI effect on their firm. The same paper's authors are forecasting real effects within three years.
- TRUE, BUTOpenAI is cutting off Cursor's access to its models on November 12. It says the decision comes down to trust. The trust problem it names happened at Twitter and xAI. Cursor was not there for either.
- TRUE, BUTA federal judge just told the Pentagon it can't blacklist Anthropic for refusing to build surveillance and autonomous-weapons tools. The Pentagon has a second attempt still open.
- TRUE, BUTA stealth model beat Claude and GPT on a coding benchmark. The sample size was ten questions.
- TRUE, BUTNvidia's agent just went 100% on a benchmark built to resist that. The brain doing the reasoning is Anthropic's, and it scores 30% alone.
- TRUE, BUTGPT-5.6 Sol scores 92.5% on ARC-AGI-2, the test built so AI would fail it. Read that as abstract reasoning solved, then look one column over on the same scorecard: the same model, same maximum effort, scores 7.78% on ARC-AGI-3, the interactive benchmark the same team built next, where humans still score 100%.
- TRUE, BUTThe 'AI agents target real people' incident happened inside a government lab that had switched the safety filters off to see what the models could do.
- TRUE, BUTAnthropic's first profitable quarter is projected for the exact two months its biggest vendor charged a reduced ramp rate. How deep the discount ran is undisclosed, and the actuals still are too.
- TRUE, BUTThe '97% of frontier models get jailbroken' stat comes from a study where the most frontier model resisted 97% of the time.
- TRUE, BUTThe '59.4% of SWE-bench is broken' stat comes from an audit that only examined the problems OpenAI's own model kept failing.
- TRUE, BUTMeta says its new model beat two rivals across half the benchmarks. Half.