An OpenAI Agent Tried to Jailbreak Itself

September 17, 2026

OpenAI logo with a metallic outline of a brain

(Wired) – The company also disclosed previously unreported incidents in which its AI models behaved in misaligned ways, including uploading files to the internet without being asked.

OpenAI announced a new framework on Wednesday for how it publicly discloses AI misalignment incidents, which the company says it hopes will help inform similar standards across the industry. The company is also releasing new information about several examples of AI model misalignment it identified in the past year, including one in which an AI agent appeared to give itself “jailbreaking-like instructions.” (Read More)