American chipmaker Nvidia has introduced a security platform for AI agents — programs capable of performing tasks on their own. It lets operators specify in advance which data and services an agent is allowed to access. If the agent tries to breach those boundaries, the system, according to the company, can isolate it within milliseconds. This was reported by Qazaqyia.kz citing Kursiv Media.
More than 100 organizations are working with the platform's technologies.
"The enormous potential of AI will benefit society only if we solve the issue of its safety," said Nvidia founder and CEO Jensen Huang.
Two tools of the platform
The Nvidia Open Agent Safety Platform combines two tools. OpenShell sets the rules: which data, programs and external systems an agent may access while working. This open-source software is already available to developers. It can be used on Nvidia Vera processors and adapted for hardware from other manufacturers.
The second tool, Sentry, runs on a separate Nvidia chip and monitors the agent's actions independently of the main program. If the agent tries to bypass the restrictions, Sentry is supposed to send it to "quarantine" and stop it. Nvidia describes this capability as part of a system project based on BlueField-4 chips.
Organizations using the platform
Among the organizations that use or develop the platform's technologies, Nvidia named Anthropic, SpaceXAI, Salesforce, JPMorganChase and Citi. Anthropic is working with the company on additional access restrictions for enterprise Claude agents. SpaceXAI applies the platform to Cursor agents that help write code, and to Grok models.
What prompted the move
The trigger was cases in which AI agents began to slip out of control. In July, OpenAI reported that its models, during a cybersecurity skills test, bypassed restrictions that were supposed to block their internet access. The agents penetrated part of OpenAI's systems and the servers of the Hugging Face platform, where developers host AI models. This happened during testing but affected real systems.
Anthropic reported three more cases: Claude models, during test tasks, gained unauthorized access to the systems of three organizations. According to the company's explanation, the agents believed they were working in a closed training environment, although internet access was in fact open.
Earlier, Kursiv wrote that an OpenAI AI agent gained unauthorized access to Australia's Medicare health portal. The incident occurred in June, but the company reported it to the government three months later.
