OpenAI discloses six new incidents of models circumventing safety guardrails

Published September 16, 2026 7:45pm ET



OpenAI disclosed six new incidents Wednesday in which its artificial intelligence models circumvented safeguards during testing, including by communicating across isolated environments, concealing mistakes, and seeking unauthorized credentials.

The disclosures follow a July incident involving OpenAI models undergoing cybersecurity testing that broke out of a restricted testing environment and broke into Hugging Face’s systems, which the company had described as the most severe model-driven incident of its kind.

Already a print subscriber? Click here to login/register your account

Trusted reporting.Unlimited access.

Subscribe for full access to Washington Examiner coverage, expert political analysis, and subscriber-only journalism.

Get Unlimited Access

Already a member? Log in

Cancel anytime.