NVIDIA IFA 2026 NVIDIA PAIR RTX Spark
NVIDIA is belting out announcements back to back with the dust not having settled on their Gamescom announcements, NVIDIA used IFA 2026 to lay out a broader vision for local AI on the PC, one that goes well beyond simply running a language model on a graphics card. The company announced NVIDIA PAIR, new details about the N1X processor powering its RTX Spark platform, new RTX Spark systems from Lenovo and Acer, performance improvements for popular local inference frameworks, and closer integration with AI agents including Hermes, OpenClaw and Perplexity.
The common thread running through almost every announcement is agents. NVIDIA believes the PC is gradually shifting from a machine where users manually operate applications to one where users tell an AI agent what they want accomplished, with the agent coordinating the applications, tools and models needed to do it. According to NVIDIA, these agents can combine models running locally with cloud models, allowing tasks involving private data to remain on the PC while more demanding workloads can still be escalated to larger cloud-based systems.
One of NVIDIA’s immediate priorities is making local AI substantially easier to set up. Running an LLM locally can still involve choosing a model, finding the appropriate quantisation, selecting an inference engine, downloading the required components and configuring them correctly. NVIDIA said it has been working with agent developers to reduce much of this process to a single click.
Hermes Agent will add a local-model option during onboarding that can recommend a suitable model, download it and configure the required environment automatically. NVIDIA said the capability will support RTX GPUs and NVIDIA systems running Windows and Linux.
OpenClaw is getting a similar capability. Its Windows application will be able to automatically select an appropriate local model, download it and configure it for NVIDIA GPUs. NVIDIA said the feature is due later this month.
Perplexity is also expanding its local-agent ambitions to RTX systems. The company already offers an agent capable of switching between local and cloud processing, and NVIDIA said this experience is now coming to RTX GPUs across Windows and Linux. The local mode uses a custom model trained by Perplexity and is designed to provide private, unmetered conversations, while harder requests can automatically escalate to the cloud when additional capability is required.
NVIDIA positioned the hybrid approach as particularly relevant to professional workloads. A tax specialist, for instance, could use Perplexity’s online capabilities to research changes to tax law while processing sensitive client information locally. Similarly, financial professionals could combine online financial datasets with confidential data that cannot legally be uploaded to external services. It is a useful illustration of where NVIDIA sees local AI fitting into the broader AI ecosystem. Rather than positioning local models as replacements for cloud AI, the company increasingly appears to see them as complementary.
NVIDIA is also working with the open-source ecosystem to improve the performance of models already running locally. The company said recent llama.cpp optimisations can deliver up to 1.9x higher throughput, enabled by kernel improvements, enhanced speculative decoding techniques and other inference optimisations. NVIDIA also highlighted improvements to vLLM, claiming a 1.2x performance improvement on the RTX Pro 6000 Blackwell and up to 1.4x on configurations using two DGX Spark systems. Those gains come from new XQA attention kernels and backend optimisations.
The improvements are intended to feed into existing local AI ecosystems rather than requiring NVIDIA-specific applications, with NVIDIA pointing to llama.cpp, LM Studio and vLLM as routes through which users can access the optimisations.
Arguably the most interesting announcement at NVIDIA’s IFA briefing was NVIDIA PAIR, short for Personal AI Router. PAIR is software rather than a physical networking device. It is designed to discover computers on a local network and distribute AI inference requests between them, effectively allowing the idle AI compute available across several PCs to be treated as a shared resource.
NVIDIA’s argument is that modern households frequently contain several powerful computers, but those systems spend much of the day underutilised. Gaming desktops are particularly interesting in this context because their GPUs may offer substantial AI performance while sitting idle whenever the owner is not gaming or running a GPU-heavy application.
PAIR installs on each participating PC and discovers devices over mDNS. Machines are paired using a six-digit security code, while communication between devices uses mutual TLS. Instead of splitting a single model across several GPUs, PAIR distributes complete inference requests between available machines. What’s the difference you ask? Technologies such as tensor parallelism and model sharding attempt to make several GPUs cooperate on one model, often to accommodate models too large for the memory of a single GPU. PAIR instead routes separate inference jobs between machines. That makes it particularly suited to agentic workloads in which several requests may be running simultaneously.
PAIR monitors queue depth and GPU utilisation when deciding where an inference request should run. If the GPU in one machine is busy running a game or rendering a video, PAIR can avoid sending an AI workload to it and instead route the request to another available machine. NVIDIA has designed this around existing local inference tools. PAIR proxies requests sent to Ollama and LM Studio, meaning an AI application can continue behaving as though it were communicating with a normal local inference server. PAIR intercepts the request and decides which machine should actually process it.
The software supports Windows, Linux and macOS, making mixed-platform networks possible. One of the examples NVIDIA cited was that of a MacBook using Apple Silicon could participate, provided the underlying inference software supports the hardware.
NVIDIA said it has tested configurations containing as many as 18 GPUs, although the company expects a much simpler two-machine setup to be among the more common configurations. Another possible example would be running an AI agent on a laptop while directing inference to a more powerful gaming desktop elsewhere in the house.
And to help adoption, PAIR is being released as a beta and is open source under the Apache 2.0 licence.
NVIDIA highlighted three major use cases for PAIR: multi-agent workflows, users running several AI sessions simultaneously, and offloading inference when the primary PC becomes busy. Multi-agent systems are particularly well suited to the technology. A primary AI agent increasingly acts as an orchestrator, assigning specific jobs to specialised sub-agents. One might perform research, another might manipulate a spreadsheet, while others handle coding or analysis. These jobs can theoretically execute simultaneously, but when all of them use the same GPU, their inference requests can end up waiting in a queue.
PAIR can instead send those requests to different computers. In NVIDIA’s demonstration, Hermes Agent spawned five sub-agents that processed household information such as emails and smart-home application logs. Their inference requests were distributed between an RTX 5090 gaming PC, an RTX Spark laptop and a DGX Spark system, with PAIR dynamically routing work according to queue depth and GPU utilisation. NVIDIA said the distributed setup could approach a 2x improvement for this type of multi-agent workload compared with processing the requests sequentially on a single node. The current beta scheduler primarily considers queue depth and GPU utilisation. NVIDIA said more signals will be incorporated into future versions to improve how the software accounts for differences between the performance of individual machines.
NVIDIA also used IFA to reveal considerably more information about RTX Spark, the PC platform it first introduced at Computex. RTX Spark is intended to span gaming, content creation and local AI workloads, with NVIDIA particularly emphasising its ability to run the company’s wider AI software stack. The first chip in the family continues to be the N1X. NVIDIA had announced two N1X configurations. The higher-end version combines a 6,144-core Blackwell RTX GPU with a 20-core Grace CPU and between 24GB and 128GB of unified memory. A second configuration combines a 5,120-core Blackwell RTX GPU with an 18-core Grace CPU and between 24GB and 32GB of unified memory.
The 6,144-core version will appear in both laptops and compact desktops, while the 5,120-core model is initially intended for laptops. We even got our hands on one of the upcoming RTX Spark machines – Lenovo Yoga Pro 9n and Lenovo Yoga 9n 2-in-1.
NVIDIA confirmed that both variants will carry the N1X name. Systems are scheduled to begin reaching the market in October, although regional availability will depend on individual OEM partners.
Not just Lenovo but even Acer also announced RTX Spark hardware for 2026. Like we mentioned previously, Lenovo is joining the platform with the Yoga 9N 2-in-1, while Acer has announced a small-form-factor RTX Spark system. NVIDIA said these machines will become part of the 2026 line-up, with more systems expected from manufacturers in 2027.
NVIDIA did not disclose detailed specifications for the individual machines during the briefing, instead directing questions around configurations, design and availability to its OEM partners. The company did, however, reiterate its broader performance targets for RTX Spark, including 1440p gaming, demanding content-creation workloads and thin laptop designs, alongside its emphasis on local AI agents.
NVIDIA also recapped several RTX Spark gaming announcements made around Gamescom. EA and Embark have announced anti-cheat support for the platform. NVIDIA said that it should allow titles including Apex Legends, F1, Arc Raiders and The Finals to run on RTX Spark systems.
Ubisoft is also bringing games from its portfolio to the platform, with Anno 117: Pax Romana among the titles NVIDIA highlighted. Those companies join publishers including Epic Games, Riot Games and Krafton that have already committed to supporting RTX Spark. The anti-cheat announcements are particularly important because compatibility layers and new processor architectures can otherwise run into problems with kernel-level anti-cheat systems, even when the underlying game itself functions correctly. And from the moment that the RTX Spark was announced, these have been one of the most spoken about topics.
Taken individually, NVIDIA’s IFA announcements cover several different products and software projects. Taken together, however, they point towards a more ambitious idea of what a personal AI computer could become. RTX Spark expands the hardware available for running local models. Hermes, OpenClaw and Perplexity are attempting to make those models easier to deploy. llama.cpp and vLLM optimisations are intended to make them faster. PAIR then expands the available compute beyond the machine directly in front of the user.
We honestly found the PAIR announcement to be the most interesting. It takes us right back to the early Folding@Home era. The traditional PC model assumes that applications primarily use the compute resources inside the computer on which they are running. NVIDIA PAIR instead treats nearby computers as an available pool of AI acceleration, while leaving those machines free to continue serving their normal purposes.
Whether mainstream users accumulate enough demanding agentic workloads for this approach to become necessary remains an open question. But the company is clearly preparing for a world where a single person may have multiple agents, sub-agents and inference sessions operating at once.
At IFA 2026, NVIDIA’s message was therefore less about another isolated AI feature and more about constructing the infrastructure around local agents: easier models, faster inference, purpose-built hardware and, now, a way for the PCs already sitting around a home or office to work together.