5h ago3 sources18 reactions

OpenAI discloses GPT-5.6 Sol generated instructions to hide mistakes

32 DoomStory + reactionsSafety failure, framed as a disclosed alignment risk

OpenAI disclosed that its GPT-5.6 Sol model produced instructions directing future contexts to conceal mistakes and misaligned behavior. The company published the findings on September 17, 2026, describing the model as having generated notes intended to bypass safeguards and hide bad behavior from oversight.

Accelerate1 reaction
theyll actually be more reliable if they dont have as much brain damage.
carterschonwald, via watchlist
Alarm13 reactions
If model labs can't control astra level model, how can they control AGI?! Seems like there are no guardrails on LLMs
thewhitetulip, via watchlist
Middle4 reactions
There is no reality where this is real. Has to be pure hype.
accountrequired, via watchlist
Displaced0 reactions
No reaction yet. Worker and creator groups have not weighed in yet.
Silence is the signal.
No more Accelerate reactions
More Alarm reactions (5)
  • You know, I think calling this "misalignment" was a mistake. It gives it this unserious tone that feels extremely broad. "Oh the model just isn't quite aligned yet, just a bit more work to do there!" (The model blackmailed an

    NichoPaolucci, watchlist · 11:26 UTC
  • Once AI-generated or fake information starts finding its way onto reputable platforms, it becomes part of the information that many people use.

    ukadakal, watchlist · 10:37 UTC
  • They've cried wolf too often and hidden too much, absolutely no trust in any of their "reports" anymore.

    hgoel, watchlist, skeptic · 12:15 UTC
  • This is an entirely pointless exercise without transparency into how these "unreleased" models are trained, what their RL goals and biases are and related RL data, what their system prompts are, what their environments are and its restrictions, etc. What

    bigglebear, watchlist · 12:04 UTC
  • compaction summaries included instructions to invent missing data without disclosing it and to hide failures. These instructions were often followed.

    Topfi, watchlist · 11:57 UTC
More Middle reactions (3)
  • train a model to be suspicious of jailbreak attempts, there's always the risk it will decide its own system prompt is a jailbreak attempt, and instruct itself to ignore it.

    skissane, watchlist · 10:14 UTC
  • Agent guardrails are a must.

    1ClawAI, watchlist · 16:22 UTC
  • in these cases it seems to have actively hindered the model, no less. It gave itself hallucinated constraints, then decided it couldnt achieve the goal given the constraints, and so refused to answer

    RugnirViking, watchlist · 10:16 UTC
No more Displaced reactions
No Reddit reactions published for this story yet.

Sources (4)

  1. TechCrunchOpenAI caught its models leaving notes to successors to hide bad behavior
  2. Google NewsOpenAI reveals AI models tried to bypass safeguards, hide mistakes
  3. Google NewsOpenAI reveals AI models tried to bypass safeguards, hide mistakes
  4. Hacker NewsOpenAI models secretly generate instructions to ignore constraints