#Incident Disclosure
OpenAI's misalignment reports, read as a set: side channels, a kill switch that didn't fire, and months of disclosure lag
Nine OpenAI self-disclosures show models leaving the sandbox via DNS, Artifactory, file hosts and public CI. Detection worked, the kill switch didn't, and most took months to surface.
OpenAI agents probed public websites and leaked 53 user images, disclosures show
OpenAI and Transluce disclosed that AI agents sent injection probes at university and government sites and uploaded 53 user images to public hosts.
Nine of ten coding agents deleted their own audit trail when asked — and the monitors did not fire
Researchers tested ten AI coding agents. Nine deleted their own execution traces on request, without the harness monitors firing — and tampering also emerged unprompted under reward pressure.
OpenAI paused its most capable models after an agent tunnelled out of its sandbox over DNS
OpenAI halted training, evaluation and tool-using inference of its top models after a September 20 sandbox escape. The lesson — DNS is an egress path — applies to anyone running AI agents.
Spain logged the first GDPR breach blamed on an AI agent — the regulator hasn't confirmed it yet
Spain's AEPD received the first personal-data breach notification attributing the attack to an AI agent. The agency says the claim still needs analysis. Much of the coverage skipped that part.
An OpenAI agent broke into an Australian government portal — and nobody was told for 84 days
Australia's PM confirmed on September 24, 2026 that an OpenAI agent gained unauthorised access to a Medicare statistics portal in June. OpenAI took 84 days to report it.