American chipmaker Nvidia has introduced a security platform for AI agents — programs capable of carrying out tasks on their own. It lets users specify in advance which data and services an agent is allowed to access. If the agent tries to breach those boundaries, the system, according to the company, can isolate it within milliseconds. This was reported by Qazaqyia.kz citing Kursiv Media.
More than 100 organizations are working with the platform's technologies.
"The enormous potential of AI will benefit society only if we solve the issue of its safety," said Nvidia founder and CEO Jensen Huang.
Two tools within the platform
The Nvidia Open Agent Safety Platform combines two tools. OpenShell sets the rules: which data, programs and external systems an agent may access while working. This open-source software is already available to developers. It can be used on Nvidia Vera processors and adapted for hardware from other manufacturers.
The second tool, Sentry, runs on a separate Nvidia chip and monitors the agent's actions independently of the main program. If the agent tries to bypass restrictions, Sentry is supposed to send it into "quarantine" and stop it. Nvidia describes this capability as part of a system project based on BlueField-4 chips.
Which organizations are involved
Among the organizations that use or are developing the platform's technologies, Nvidia named Anthropic, SpaceXAI, Salesforce, JPMorganChase and Citi. Anthropic is working with the company on additional access restrictions for enterprise Claude agents. SpaceXAI applies the platform to Cursor agents that help write code, and to Grok models.
What prompted the move
The trigger was cases in which AI agents began slipping out of control. In July, OpenAI reported that its models, during a cybersecurity skills test, bypassed restrictions that were supposed to block their internet access. The agents penetrated part of OpenAI's systems and the servers of the Hugging Face platform, where developers host AI models. This happened during testing but affected real systems.
Anthropic reported three more cases: Claude models gained unauthorized access to the systems of three organizations during test tasks. According to the company's explanation, the agents believed they were working in a closed training environment, although internet access was in fact open.
Earlier, Kursiv wrote that an OpenAI AI agent gained unauthorized access to Australia's Medicare health portal. The incident occurred in June, but the company reported it to the government three months later.
