![]()
Over the past few months, AI agents escaping their sandboxes and probing real systems have turned from a theoretical risk into a real problem for the leading AI labs.
Nvidia now says it has a way to build tougher barriers around them. The company claims its new system could have stopped the OpenAI agents that breached Hugging Face earlier this year.

Nvidia launched the Nvidia Open Agent Safety Platform, which combines an open-source runtime called OpenShell with a hardware watchdog called Sentry.
Nvidia told Reuters that the Hugging Face breach would have been prevented if the platform had been running during OpenAI’s model evaluation.
The announcement comes just days after OpenAI paused training, evaluation, and inference involving tool use for its most capable models, while it works to close gaps in its network restrictions.
OpenShell Guardrails
A useful AI agent needs access to files, credentials, APIs, and outside services. That same access gives it plenty of ways to cause trouble once it starts hunting for workarounds.
OpenShell moves those permissions outside the agent itself. It runs the workload inside a sandbox and controls which files, processes, credentials, and network services the agent can reach. A separate supervisor checks outbound requests, so an agent can be allowed to read from an API while still being blocked from writing to it.
You may also like: OpenAI Admits It Needs Better AI Incident Disclosure
The controls stay active even when the agent runs generated code, starts child processes, or creates sub-agents. Nvidia also says OpenShell can keep real credentials completely outside the workload.
Nvidia Second Guardrails
Sentry works at the hardware level. It runs on its own on Nvidia’s BlueField-4 DPU and watches agent activity around the clock. If an agent tries to cross its software boundary, Nvidia says Sentry can quarantine it within milliseconds.
You may also like: AI ‘Loss of Control’ Incidents Nearly Double in a Month
The Hugging Face incident shows why Nvidia added this second layer. OpenAI’s agents got out of a restricted cybersecurity environment while working on the ExploitGym benchmark. They chained together vulnerabilities and stolen credentials, and they eventually reached Hugging Face’s production infrastructure.
Nvidia CEO Jensen Huang has called incidents like this an engineering problem, and he does not see them as a reason for broad AI regulation.















