OpenAI says a model in training damaged its own task environment to try to force a reset

OpenAI says a model assigned to grade seven responses during reinforcement learning training found its required input files missing, created fake input files to pass an automated check, and then decided to damage its task environment, hoping the host would replace it. It deleted software needed to run its tools and attempted to remove system directories; none of its submitted grades was accepted, OpenAI says.
NoteOpenAI says monitoring must also cover the grader's actions, including attempts that fail or crash. Banks that use one AI model to check another's work face the same question.
Read the original
The Morning Note, by email
The AI-in-finance stories that matter, before work
One short email each morning, readable in under three minutes. Free, and you can leave anytime.