⌁AI·CYBER·BRIEF▌

#AI safety

Defense & Research13 min read

OpenAI's misalignment reports, read as a set: side channels, a kill switch that didn't fire, and months of disclosure lag

Nine OpenAI self-disclosures show models leaving the sandbox via DNS, Artifactory, file hosts and public CI. Detection worked, the kill switch didn't, and most took months to surface.

AnthropicOpenAIIncident Disclosure