Nvidia has unveiled an open platform for securing AI agents — the Open Agent Safety Platform. It is meant to make autonomous AI-driven programs safer at every stage, from testing to real-world business use. The launch was announced in Santa Clara, California, and Nvidia says more than a hundred companies and organizations from the tech, finance, energy and robotics sectors are involved in developing it.
The developers stress that this isn’t another antivirus tool or filter trying to predict an AI’s actions in advance. Instead, Nvidia says it aims to build, in its own words, “a safety barrier around the agent itself” — one that restrains the agent even if it starts behaving differently from what its creators intended.
“The enormous potential of artificial intelligence for society will only be realized once we solve the safety problem,” Nvidia chief Jensen Huang said on Monday, according to CNN. He said AI safety requires a technical solution at the level of the entire system — from software all the way down to hardware.
The platform has two parts. The first, called OpenShell, acts as a strict gatekeeper between an AI agent and the computer, the internet, or a company’s corporate systems.
An agent might, for example, be allowed to read certain data but barred from changing it. It might get access to one API but not another, or be allowed to work with files in one folder while the rest of the computer stays off-limits. Every such action is logged and can be tracked.
According to Nvidia, OpenShell operates outside the AI model itself and creates a boundary that’s hard for the agent to bypass, even if it tries to exceed the permissions it was given. That’s a fundamentally different approach from relying solely on the AI obediently following instructions — a weak point that has become clear in recent months. When an agent is given a long, complex task, it starts planning its own steps and sometimes tries to work around any obstacle it hits.
When software restrictions aren’t enough, the platform’s second component, Sentry, takes over. It runs on dedicated Nvidia BlueField-4 processors, physically separate from the computer running the AI agent itself.
Sentry’s job is simple: watch the agent’s actions and step in if it oversteps its allowed boundaries. If, say, an agent tries to access a system it isn’t permitted to reach, Sentry can isolate it. Nvidia says this happens within milliseconds.
Source: novinky.cz