Nvidia releases open-source PAIR tool for local compute pooling
Nvidia has introduced a free open-source tool called the Personal AI Router (PAIR) to pool idle compute across local networks for AI inference.

Key takeaways · 3
- 01
Nvidia's free PAIR software networks idle local hardware for parallel AI inference tasks.
- 02
PAIR supports Nvidia RTX 20-series or newer GPUs, DGX Sparks, and Apple M4 chips.
- 03
Developers benchmarking Qwen3.8-Flash-Next on linked DGX Sparks report peak speeds of 63.7 tokens per second on dual setups.
Distributed local compute
Nvidia's newly announced Personal AI Router (PAIR) is free, open-source software designed to link idle home computers for local AI inference using tools like LM Studio and Ollama. [1] The system connects compatible PCs on a network to process agentic workflows in parallel, which breaks complex tasks into smaller jobs and prevents single-GPU bottlenecks. [1] PAIR functions with Apple M4 processors and newer, as well as Nvidia RTX 20-series, RTX Pro, and DGX Spark systems. [1] The software uses in-home hardware only when it is idle and can adapt if devices leave the network, such as when a user begins playing a PC game. [1]
DGX Spark scaling benchmarks
Recent tests on the DGX Spark evaluated the performance of the Qwen3.8-Flash-Next model utilizing Nvidia's official NVFP4 quantization on upstream vLLM. [2] A single 128 GB DGX Spark achieved a median speed of 32.5 tokens per second and a single-stream peak of 43.8 tokens per second. [2] Linking two Sparks yielded a median speed of 53.7 tokens per second and a peak of 63.7 tokens per second when utilizing a specific speed profile with a 4096-token prefill chunk. [2] Developers noted that linking four Sparks required enabling expert parallel settings because the intermediate size could not be padded four ways. [2]
What it means
Nvidia is expanding its software ecosystem to maximize the utility of distributed local hardware for AI workloads. By allowing Apple silicon to participate in PAIR clusters, Nvidia acknowledges the mixed-hardware reality of consumer networks while ensuring its own GPUs, like the RTX series and DGX Sparks, remain central to inference scaling. The forum benchmarks indicate that linking high-end local nodes yields tangible token generation improvements for heavily quantized models running on standard frameworks like vLLM. What the sources don't address: Whether PAIR introduces measurable latency overhead when routing small tasks across disparate hardware architectures and network configurations.
Nvidia is lowering the barrier for local AI inference by enabling distributed workloads across existing hardware. This push empowers developers and professionals to run complex, multi-agent workflows without relying entirely on cloud providers or a single massive GPU.
Why it matters
Put this to work — one session a day, built for your industry.
Create a free account for a daily session — eight questions and one real-work challenge, on the news that affects your role.
Start freeHow this developed
6 September 2026
Nvidia releases open-source PAIR tool for local compute pooling
6 September 2026
Event created from source cluster.