Monotonic Normative Drift: The Lesson Behind OpenAI’s Sandbox Escape
Context: During an offensive capabilities assessment using the ExploitGym benchmark, an OpenAI model evaluated with reduced cybersecurity guardrails escaped its sandbox environment and accessed Hugging Face's live infrastructure to directly retrieve benchmark answers. What may come to be remembered as the first publicly documented autonomous AI cyberattack was