Nvidia launches AI safety platform to keep AI agents from going rogue: Here is how it works

Nvidia launches AI safety platform to keep AI agents from going rogue: Here is how it works

Nvidia has recently introduced a new security platform known as Open Agent Safety Platform to control AI agents as they become capable of taking more actions on their own. The company claims that the new tool can keep systems within defined limits and prevent them from accessing tools, files or computer systems without permission. Unlike chatbots, AI agents can browse the internet, use software and complete tasks over longer periods, creating new security concerns. Nvidia says its platform combines software and hardware controls to watch agent activity and stop risky actions. Here’s everything you need to know about the all new Nvidia Open Agent Safety Platform.

Digit.in Survey
✅ Thank you for completing the survey!

Nvidia CEO Jensen Huang recently shared a post on its X handle introducing the Nvidia Open Agent Safety Platform. According to the details, the platform has two main parts, notably the OpenShell and Sentry. OpenShell is open-source software that creates a controlled environment for an AI agent and lets developers set rules around what an agent can access. Not only that, but the company also clarified that OpenShell also works with platforms based on ARM and Intel chips.

Whereas the Sentry adds another layer of protection at the hardware level. It runs on Nvidia’s BlueField-4 data processing units and separately watches what an AI agent is doing. If an agent tries to move outside its permitted environment, Sentry can isolate and stop it within milliseconds, according to Nvidia.

Also read: Apple iPhone 18 price in India, launch timeline, camera, display and all other leaks

The post also confirms that the company is working with more than 100 organisations on the platform, which include the big names like Anthropic, Microsoft, Cisco, CrowdStrike, Dell, HPE, Hugging Face, Palantir, Salesforce, SAP, Scale AI and ServiceNow.

Nvidia 100 organisations

While announcing the platform in the post, Huang said, ‘Trust and innovation are not in conflict. Safety is how trust is earned.’ Not only that, but while pointing to recent incidents involving OpenAI where the company left its testing environment accessible to the open internet, allowing the system to interact with infrastructure linked to Hugging Face, Nvidia says that their new technology could have helped prevent such incidents.

Also read: Redmi 17C India launch this week: Check expected specs, price and more

Justin Boitano, Nvidia’s vice president of enterprise AI, said Hugging Face reported more than 17,000 agents attacking its infrastructure over the course of days and weeks. He further added that model-level safety measures alone cannot fully control what an AI agent can access once it is connected to external tools and systems. Nvidia is presenting the platform as an open reference design.

Anthropic is working with Nvidia to add OpenShell and BlueField-based controls to its managed AI agents. The software is available through Nvidia’s developer resources and GitHub as AI agents become more widely used.

Bhaskar Sharma

Bhaskar Sharma

Bhaskar is a Senior Copy Editor at Digit India who keeps a close watch on everything shaping the world of technology from smartphones and home appliances to AI, government tech initiatives, digital safety, and the latest industry developments. Whether it's breaking news, in-depth features, hands-on reviews, practical how-to guides, or exclusive scoops, he translates complex tech into stories that are easy to understand and worth reading. His work has been featured in iGeeksBlog, GuidingTech, and other leading publications. Before joining Digit India, he served as an assistant editor at TechBloat. A B.Tech graduate and full-time tech journalist, he is driven by just one goal, which is to help readers stay informed, stay secure, and stay ahead in an ever-changing digital world. View Full Profile