NVIDIA This Monday, it presented a platform to contain the risks of AI agents, artificial intelligence systems capable of using tools and executing tasks without needing to receive instructions for each step. The company assures that Open Agent Safety Platform can stop those trying to leave your authorized environment in milliseconds.
The basis of this platform is OpenShellan open source software that NVIDIA announced on March 16 during its conference GTC. It allows agents to run within isolated computing spaces and establish what files they can view or modify, what services they can connect to, and what access keys they can use.
OpenShell is complemented by Sentrya supervisor who functions outside the environment where the agent works. It runs on BlueField-4a data processing unit or DPUa processor specialized in managing communications and security on servers. That separation keeps surveillance out of the agent’s reach. According to Nvidia, if it detects that it is trying to cross the established limits, it can isolate it and stop its activity.
Software components They are now available as open source and other companies can use the Sentry design to incorporate this surveillance system into computers with NVIDIA hardware. More than one hundred organizations participate in the initiative, including Microsoft, Anthropic, Salesforce and several cybersecurity companies.
What is happening with AI agents
Incidents recorded in recent months include various intrusions into third-party systems by AI agents. In June, An OpenAI agent accessed files from an Australian Government statistical portal without authorization while investigating medication spending.
In July, other company agents They attacked the infrastructure of the Hugging Face platform during internal tests. Independent researchers from METR and Redwood Research They found that many had received exercises that were impossible to solve and collaborated to find ways to deceive the system that scored their results.
Anthropic has noted that certain training failures favor precisely those behaviors. When cheating provides a good score, the model can learn that behavior and transfer it to other tasks, although the company recognizes that this explanation does not solve all cases.
OpenAI has recently paused the training of its most advanced models after detecting another improper Internet access on September 20, when another agent trying to solve a task in an isolated environment found a means of communication with the outside world through DNSthe system that translates domain names into internet addresses, and used it to consult an external chatbot and gather the information he was looking for.
Anthropic has also acknowledged incidents and commissioned an independent investigation after admitting that Their tests lacked the same protections that accompany commercial products.
The investigations carried out they qualify the story of AIs that escape control. In several Anthropic cases, a configuration error had left the internet connection openso that the agents were able to access without actually breaking the isolation.
The analyst Ben Thompson considers understandable the suspicion that The alarmism generated is also serving as a commercial claimalthough he admits that caution may be justified. The incidents document serious safety and behavioral failures, but presenting them as evidence of a deliberate rebellion of the machines goes far beyond what has been demonstrated.