Microsoft and NVIDIA’s new RTX Spark laptops cost Rs 300000, and they can run 120B+ parameter models locally

HIGHLIGHTS

Microsoft is positioning Windows as a hybrid AI platform where workloads can move between local hardware and the cloud.

NVIDIA RTX Spark PCs can offer up to 128 GB of unified memory and run AI models exceeding 120 billion parameters locally.

Microsoft Execution Containers aim to give autonomous AI agents tighter controls over files, networks and system access.

The AI PC has spent the past couple of years largely being defined by NPUs and the TOPS that can be churned out by the CPU, GPU and NPU put together. That’s a lot of marketing speak but not as much of a viable use case. That has changed for the better in the recent months with the advent of smarter open source AI models and easier-to-use AI tools. With that being a sign of the changing times, Microsoft and NVIDIA now appear to be setting their sights considerably higher. 

At its latest Windows AI and Surface event, Microsoft unveiled the Surface Laptop Ultra and laid out what it calls its vision for hybrid intelligence, where AI workloads are no longer tied exclusively to either a PC or the cloud. Instead, Windows will increasingly decide where a workload should run based on factors such as compute requirements, available hardware and the capabilities required for a particular task. NVIDIA, meanwhile, is supplying much of the hardware muscle behind the more ambitious end of that vision through its new RTX Spark platform that was first showcased at COMPUTEX 2026. Together, the two companies are effectively proposing a different future for the Windows PC: one where sufficiently powerful laptops and desktops can run large AI models, host persistent agents and perform workloads that would previously have been sent to cloud infrastructure. The Surface Laptop Ultra is priced at USD 2599 and could see prices starting at INR 300000 in India.

Microsoft calls its approach ‘hybrid intelligence’

Microsoft’s central argument is that AI does not necessarily have to choose between local and cloud computing. Under its hybrid intelligence paradigm, Windows can use models running directly on the PC when appropriate while still reaching into cloud-based models when more compute or different capabilities are required. The idea is partly about performance and privacy, but Microsoft is also positioning local inference as a way of controlling the rising cost of cloud AI. It’s not a new concept, routers have existed for a while now and services such as Perplexity and OpenAI have had them for quite a while wherein the choice of AI model to be used is determined on the basis of the complexity of the prompt. Now that concept is coming to your PC. 

The company says customers increasingly want access to sophisticated models without consuming cloud tokens for every task. Local models can therefore take care of workloads where the PC has sufficient compute, while cloud models remain available when required. GitHub is one of the first examples of how this may work. Microsoft says GitHub’s HydraFusion model-routing technology is being extended to Windows so it can route requests not only between cloud models but also models running directly on a user’s machine. Support is expected to arrive experimentally in GitHub Copilot, GitHub Copilot CLI and Visual Studio Code later in October when GitHub is expected to unveil the feature at GitHub Universe, their annual flagship event.

This is potentially a more meaningful evolution of an AI PC that’s a lot more utilitarian. It’s no longer about throwing in an NPU and calling something an AI PC. With an improved synergy between the hardware and the software bits, PC becomes one part of a distributed AI infrastructure, with workloads able to move between CPU, GPU, NPU and cloud resources.

Large AI models are coming directly to Windows PCs

The scale of the models Microsoft expects these systems to run is also changing. Microsoft says its MAI Code 1.1 Flash model contains 137 billion total parameters, with 6.8 billion active parameters. Through 3-bit quantisation, the company says it has reduced the model’s size by almost 80 per cent while retaining support for a 256K-token local context window. That’s a huge improvement considering that just two years ago, running such a large model would have required you to pack oodles of RAM on your laptop aside from a powerful NPU. 

Windows will also support an upcoming NVIDIA Nemotron model with more than 70 billion parameters, compressed using 2-bit quantisation to consume a little over 20 GB of memory. Microsoft also cited DeepSeek V4 Flash, a 284-billion-parameter model, among the models being targeted for local execution on RTX Spark hardware.

Microsoft is simultaneously adding llama.cpp support to Windows ML, giving developers a more direct route to running open-source models using Microsoft’s AI runtime across GPUs, NPUs and CPUs. The common thread is quantisation and large pools of memory. Instead of attempting to run full-precision frontier models, increasingly aggressive compression allows large models to operate within the memory limits of high-end personal computers.

NVIDIA RTX Spark brings the brawn

That brings NVIDIA RTX Spark laptops into the picture. The platform combines a Blackwell RTX GPU with as many as 6,144 CUDA cores and an NVIDIA Grace CPU with up to 20 cores. The CPU and GPU are connected through a 600GB/s interface, while systems can be configured with as much as 128 GB of unified memory. Think of it as the DGX Spark compute mini-super computer but for Windows instead of Linux. And just like the DGX Spark, NVIDIA rates the RTX Spark platform for up to one petaflop of FP4 AI performance. 

Unified memory is one of the key differentiators. Large language models are extremely memory intensive, and having a large shared pool available to the CPU and GPU makes it possible to load models that simply would not fit within the dedicated VRAM available on conventional notebook GPUs. Microsoft says its new Surface Laptop Ultra laptop, which was unveiled during the showcase,  can consequently run models exceeding 120 billion parameters locally when equipped with 128GB of unified memory.

NVIDIA is also pitching RTX Spark as more than an AI accelerator. The architecture supports CUDA, fifth-generation Tensor Cores, NVFP4 AI processing, AV1 and 4:2:2 video encode and decode, alongside familiar RTX technologies including ray tracing and DLSS. The company claims RTX Spark systems can run AAA games at more than 100fps at 1440p using technologies including DLSS 5, Reflex and G-SYNC, although actual performance will naturally vary significantly depending on the title and system configuration. During the showcase, Microsoft showcased Gears of War running at more than 60 FPS on their Surface Laptop Ultra.

RTX Spark is becoming an entire PC category

Microsoft and NVIDIA are not limiting the platform to Surface. We’ve previously spoken about the multitude of brands which had unveiled their own versions of the RTX Spark systems. Among the current list of partners are ASUS, Dell, HP, Lenovo, MSI and Microsoft, with NVIDIA also naming Acer and Gigabyte among its hardware partners. Designs span premium notebooks, compact desktops and development systems. Microsoft has announced the Surface Laptop Ultra alongside systems including the ASUS ProArt P16 and P14, Dell XPS 16 Creator Edition, HP OmniBook Ultra 16, Lenovo Yoga 9n 2-in-1 and MSI Prestige N16 Flip AI+.

As of today, RTX Spark laptop pre-orders have opened, with systems scheduled to begin shipping from 16 October. Compact desktop systems are expected to follow later in the year. Microsoft is making some sizeable performance claims as well. Against an Apple MacBook Pro 16-inch with M5 Pro and 64 GB of memory, it claims tested RTX Spark systems can deliver up to 2.1x faster time to first token, 4.3x faster AI image generation and 6.2x faster AI video generation. Of course, real benchmarks are just a week away as the units will soon start shipping next week and scores of users are expected to put RTX laptops through the grind.

Microsoft Execution Containers (MXC)

Running AI locally is only one part of Microsoft’s strategy. The company also expects AI agents to gain substantially more access to the PC itself. These agents could potentially read files, execute code, access networks, manipulate applications and continue working in the background. That requires a different security model from conventional applications. Microsoft’s answer is Microsoft Execution Containers, or MXC, which is now generally available on Windows 11. MXC allows organisations to specify which files and networks an agent can access, with Windows enforcing those restrictions while the agent is operating. The platform can use several forms of isolation, including process and session isolation, virtual machines and Windows Subsystem for Linux containers.

Microsoft is essentially trying to give agents three things: containment, their own identifiable activity and centralised management. So, if an autonomous agent changes a file or accesses a resource, organisations need to know whether the action came from the human user or an AI acting on that person’s behalf. Microsoft says Codex, GitHub Copilot, OpenClaw, Replit, LM Studio, NVIDIA OpenShell and Unsloth AI already support MXC, while support is planned from several other platforms including Anthropic Claude Code, Perplexity and Manus. Meta’s upcoming Muse for Windows agent will also integrate with MXC.

Copilot is becoming increasingly local

Microsoft’s own Copilot is also being rebuilt around this hybrid model. On Copilot+ PCs, Microsoft says Copilot will eventually gain three important abilities: access to local context, the ability to perform local actions and access to models running directly on the computer. With permission, Copilot could understand files and recent activity on the PC. It could also perform actions such as organising files, checking device diagnostics, troubleshooting problems and carrying out workflows. Tasks could then be assigned to either local or cloud models depending on what makes sense. 

This also extends Microsoft’s idea of Copilot beyond a chatbot sitting inside an application. The company’s Autopilot experience is intended to operate as a persistent agent capable of continuing work while the user focuses elsewhere. Microsoft says these hybrid intelligence features will begin rolling out to Copilot+ PCs over the coming months, although availability will differ according to hardware, market and silicon platform.

And then there is DGX Station for Windows

RTX Spark may represent the laptop and compact PC tier, but Microsoft and NVIDIA are going considerably further at the high end. NVIDIA has previewed DGX Station for Windows, bringing its GB300 Grace Blackwell Ultra Desktop Superchip to a Windows-based workstation. The machine offers 748 GB of coherent memory and up to 20 petaflops of FP4 AI compute. Microsoft says that is sufficient to run models at around the one-trillion-parameter scale locally. 

This also marks a notable change for NVIDIA’s DGX Station products, which have traditionally operated within Linux environments. Windows support means organisations could theoretically run large-scale local inference and agentic workloads on the same platform used for their conventional enterprise applications, while Linux development environments remain accessible through WSL. Microsoft also imagines these machines functioning as shared “token factories”, with a single DGX Station capable of serving 32 or more simultaneous AI agents for developers, researchers and engineers.

The AI PC is starting to become something different

The significance of these announcements is less about any individual laptop and more about how Microsoft is redefining the role of the PC in AI computing. The first wave of AI PCs primarily demonstrated that selected AI features could run without the cloud. Microsoft’s latest vision asks a much bigger question: how much of the AI infrastructure currently sitting inside data centres can eventually move back onto personal and enterprise computers? RTX Spark systems with 128 GB of unified memory are still likely to sit towards the expensive end of the PC market, while DGX Station clearly targets specialised professional and enterprise workloads. But the underlying direction is notable.

Microsoft is building Windows around a world in which local AI models, cloud models and autonomous agents coexist. NVIDIA is supplying hardware powerful enough to make increasingly large models practical on the desktop, while Windows is being equipped with the routing, runtime and security layers needed to manage them. If that model works, the “AI PC” may stop being defined by individual AI features. Instead, the PC itself becomes a local AI compute platform, capable of running models and agents continuously while calling on the cloud only when it actually needs to.

Mithun Mohandas

Mithun Mohandas is an Indian technology journalist with 14 years of experience covering consumer technology. He is currently employed at Digit in the capacity of a Managing Editor. Mithun has a background in Computer Engineering and was an active member of the IEEE during his college days. He has a penchant for digging deep into unravelling what makes a device tick. If there's a transistor in it, Mithun's probably going to rip it apart till he finds it. At Digit, he covers processors, graphics cards, storage media, displays and networking devices aside from anything developer related. As an avid PC gamer, he prefers RTS and FPS titles, and can be quite competitive in a race to the finish line. He only gets consoles for the exclusives. He can be seen playing Valorant, World of Tanks, HITMAN and the occasional Age of Empires or being the voice behind hundreds of Digit videos.

Connect On :