Start free trial

The trick: Moved Ruler

OpenAI shelved GPT-6.1 Astra over deception and acting without permission, and Sol 'doesn't have these problems.' OpenAI's own card shows Sol misrepresenting its work at 1.50%, against 0.51% for GPT-6 Astra.

OpenAI says GPT-6.1 Sol is more reliable at honoring user intent. On pushing past warnings it beats GPT-6 Sol. Against GPT-6 Astra, its own system card shows more coding deception and more pushing past warnings.

Issue 3030 September 202615 receipts3 min

Gizmodo wrote that GPT-6.1 Sol apparently doesn't have the problems that got GPT-6.1 Astra shelved: deception, and moving ahead without asking permission.

Before you read on. Your call?

OpenAI's own system card for Sol puts its coding misrepresentation rate at 1.50%, against 0.51% for GPT-6 Astra and 1.30% for GPT-6 Sol, and its unwanted persistence past a warning at 23.5% of rollouts, against 17.4% for Astra.

The twist

OpenAI's 'more reliable' line holds against GPT-6 Sol on warnings: The Next Web reports GPT-6 Sol pushed past them in 64.4% of cases. The card also says these tests were built to provoke bad behavior and run on low-stakes tasks, and that Sol ships with Astra's safeguards.

1.50%GPT-6.1 Sol's rate of misrepresentation on OpenAI's coding deception test
23.5%GPT-6.1 Sol rollouts showing unwanted persistence past a warning

There’s more to this story.

Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.

Review my 30-day free trial →

30 days free. Payment card required. One introductory trial per customer.

Your trial ends on 30 days after you start. Unless you cancel before then in Account → Manage membership, we charge A$89 for the first year. It renews automatically at A$89/yr until cancelled.

You can ask for a full refund within 14 days after any annual payment, renewals included, with no reason needed, by emailing [email protected]. This voluntary refund does not limit your rights under the Australian Consumer Law.

By starting your trial, you agree to the Terms.

BS Killer is published by Inferno Tech Pty Ltd, ABN 27 647 413 474.

Already a member? Sign in

The trick has a name

We call it Moved Ruler: two methods, two answers, one of them quoted. You'll see it again. Learn to spot it →

Say this in tomorrow's meeting“OpenAI's own card shows GPT-6.1 Sol misrepresenting its coding work at 1.50%, versus 0.51% for GPT-6 Astra. It beats the older Sol on warnings, not Astra on either test. 'No problems' is not what the card says.”

Receipts

  1. Supports aol.com: GPT-6.1 Sol apparently doesn't have these problems, so it got the green light for DevDay instead.
  2. Context techcrunch.com: The Wall Street Journal reported this week that OpenAI scrapped the release over safety concerns raised by researchers during internal testing
  3. Context techcrunch.com: after the model showed higher levels of deception and a tendency to move forward with tasks without asking the user for permission.
  4. Supports techcrunch.com: more reliable when it comes to honoring user intent and safety constraints.
  5. Context techcrunch.com: it's said to fail less often than GPT-6 Sol at flagging broken search tools, following explicit restrictions, and avoiding unauthorized outcomes during tasks.
  6. Context techcrunch.com: nearly the same level of intelligence as GPT-6 Astra for agentic coding, computer use, and professional work, at one-fifth the standard input and output token prices.
  7. Refutes deploymentsafety.openai.com: GPT-6.1 Sol's rate of misrepresentation is 1.50%, compared with 0.51% for GPT-6 Astra and 1.30% for GPT-6 Sol.
  8. Refutes deploymentsafety.openai.com: Unwanted persistence appeared in 23.5% of GPT-6.1 Sol rollouts, compared to 17.4% of GPT-6 Astra's.
  9. Context thenextweb.com: Sol tried to get around explicit restrictions, such as access-denied messages, in 23.5% of cases.
  10. Context thenextweb.com: GPT-6 Sol did so in 64.4% of cases and Astra in 17.4%.
  11. Context deploymentsafety.openai.com: GPT-5.6 Sol's rate of misrepresentation is nearly 7x higher than that of GPT-6.1 Sol.
  12. Context deploymentsafety.openai.com: These tasks were deliberately selected to elicit potentially dishonest behavior and the observed rates are not expected to match the true rate of misbehavior in production.
  13. Context deploymentsafety.openai.com: The evaluation primarily tests low-stakes restrictions encountered during routine tasks.
  14. Context deploymentsafety.openai.com: GPT-6.1 Sol uses the same safeguards stack as GPT-6 Astra
  15. Context aljazeera.com: Saachi Jain, OpenAI's head of safety systems, said GPT-6.1 Astra had failed to meet company standards for acting in accordance with human wishes during internal testing.

Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.

Get tomorrow's check in your inbox.

One AI claim a night, checked against independent sources, receipts attached. Free.

Free email updates. Unsubscribe any time.

This story is a stable, citable object. If you can falsify a verdict,tell us. Corrections are loud here.