Ever since ChatGPT came out and captivated the whole world roughly three years ago, AI’s biggest flex has been all about scale. Bigger models needed bigger clusters of datacentres for all their bigger and better training runs. NVIDIA, obviously, benefited from all of this, and they have the stock price to prove it!
However, NVIDIA’s latest announcements suggest the next phase of AI growth could be headed in a different direction. It seems the question is no longer about just how large an AI model can become, what’s going to matter more is how quickly, cheaply and locally AI’s benefits can be delivered on-device at the same scale NVIDIA has helped architect in the cloud.
At one end of the AI spectrum sits Groq 3 LPX, which is now finally in full production for NVIDIA’s Vera Rubin platform. According to NVIDIA, the Groq 3 LPX chip can churn out 3,400 output tokens per second running the 31-billion-parameter Gemma 4 model with a 100,000-token context window. And while it does all of that, the new NVIDIA chip also delivers four times faster responsiveness than the next best chip alternative.
So there’s Groq 3 LPX on one end of the stick, geared towards more agentic AI workflows in the datacentre, but at the same time on the other end NVIDIA also revealed the updated Jetson Orin Nano 2. It’s a compact 78-TOPS robotics computer with 8GB dedicated memory, and NVIDIA claims it’s twice as fast in overall inference performance compared to the previous generation, while also being more power efficient overall – the figure they’re sharing is Jetson Orin Nano 2 consumes 40% less power at equivalent performance than the previous generation.
So yeah, these are different products, but with them NVIDIA is placing the same overall bet, if you read on.
Training created the generative-AI boom, it’s what allowed everything from ChatGPT to Gemini to become capable of being able to do intelligent tasks. Those tasks are where inference comes into the picture, where all the AI models get to work to answer a question, generate code or start an agentic AI workflow. And inference is the new AI battleground.
Inference speed optimization matters for agentic AI, and NVIDIA knows it better than anyone else. Unlike a chatbot that only responds to one question at a time, an AI agent is always doing more than one thing – it may inspect files, call additional tools, generate code, test its accuracy and repeat the loop hundreds of times or even more. NVIDIA’s Groq 3 LPX chip is built specifically to cater to agentic AI, where token-generation latency matters more than anything else.
According to an IDC report, worldwide AI-infrastructure spending reached $318 billion in 2025, more than double 2024’s $153 billion, and more will be spent through 2026 – $487 billion to be exact. However, only about 11% companies are implementing agentic AI use cases, according to a Delloite report in 2026. Faster inference is one of the key requirements for that adoption number to increase through 2026-2027, where demos can transition into dependable products.
While Groq 3 LPX is still aimed inside datacentres and AI factories, Jetson Orin Nano 2 will be important for consumer and end user AI use cases. NVIDIA’s positioning the Jetson Orin Nano 2 class of chips for drones, home robots and vision systems. In a briefing call, NVIDIA said there’s interest from partners exploring delivery drones and domestic cleaning robots that can talk to humans.
Stanford’s AI Index found that the SLMs clearing 60 percent capability of LLMs fell from 540 billion parameters in 2022 to 3.8 billion parameters in 2024, which is a whopping 142x reduction. When model sizes fall, so do their cost, suggests Stanford, where the cost of a GPT-3.5-level inference reduced by 280x between late 2022 and late 2024 period.
What this means is simply that not only model sizes are shrinking, that size reduction is allowing them to be deployed on-device without a lot of dependence on the cloud. The overall global edge-computing spend is expected to reach $450 billion by 2029, according to IDC. This will directly mean AI infusion in more and more consumer-end devices – not just smartphones and laptops, but appliances, robots, cars. That’s the opportunity Jetson Orin Nano 2 is attacking.
Looking at both the NVIDIA announcements, it feels like the bigger takeaway is more architectural in nature. Leading tech companies will still need enormous AI factories for bigger, more complex models, and critical infrastructure tasks.
But at the same time, smaller, leaner models will start occupying more space inside physical devices that end users actually touch and feel. And it’s pretty clear that NVIDIA wants a seat at both these tables, where Vera Rubin plus Groq 3 LPX remains first choice for anyone desiring hyperscale inference, while Jetson takes the shape of embodied edge AI. Of course, the two ends of the table will be tied together by NVIDIA’s own burgeoning software ecosystem.
Jetson Orin Nano 2 is expected in the first half of 2027, and NVIDIA has not announced pricing in the release. For consumers, this is where “AI everywhere” starts to feel like it’s strategically overhauling modern tech infrastructure. AI chatbots was only the first interface, the next wave is all about whether agents respond instantly, robots and drones map and exist in their surroundings more reliably, and devices can draw AI inference without constantly pinging servers in a faraway datacentre.