All AI News
    NVIDIA BlogThursday, September 3, 2026 9 min read
    AI

    Sparks Fly: NVIDIA Accelerates Local AI at IFA 2026

    NVIDIA is collapsing the gap between cloud AI capability and local hardware, making agent deployment nearly friction-free.

    Key takeaways
    • 01Local AI just got a serious infrastructure upgrade.
    • 02NVIDIA's IFA 2026 push bundles inference speed gains up to 1.9x, a network-distribution tool called PAIR, and simplified agent setup across Hermes, OpenClaw, and Perplexity Portable Computer.
    • 03October brings RTX Spark mini-PCs from Lenovo and Acer.
    • 04August's model wave—Nemotron 3.5, Qwen3.8, Meta Muse Glimmer, DeepSeek v4 Flash—signals that frontier-class models are increasingly designed with local NVIDIA hardware as a first-class target.
    Koko brief

    NVIDIA is collapsing the gap between cloud AI capability and local hardware, making agent deployment nearly friction-free.

    Local AI just got a serious infrastructure upgrade. NVIDIA's IFA 2026 push bundles inference speed gains up to 1.9x, a network-distribution tool called PAIR, and simplified agent setup across Hermes, OpenClaw, and Perplexity Portable Computer. October brings RTX Spark mini-PCs from Lenovo and Acer. August's model wave—Nemotron 3.5, Qwen3.8, Meta Muse Glimmer, DeepSeek v4 Flash—signals that frontier-class models are increasingly designed with local NVIDIA hardware as a first-class target. **Watch:** Whether NVIDIA PAIR's cross-PC inference routing becomes a developer-adopted standard or a niche curiosity.

    Watch: Whether NVIDIA PAIR's cross-PC inference routing becomes a developer-adopted standard or a niche curiosity.

    In brief · from blogs.nvidia.com

    Frontier intelligence is going local. At IFA 2026, NVIDIA, Microsoft and its partners are teaming up to provide faster inference and new tools that make agents easier to set up and run locally on NVIDIA hardware. New compact NVIDIA RTX Spark Windows PCs are also coming in October to give AI enthusiasts, developers and creators more ways to run capable agents locally and securely. Today’s announcements include: Simplified local AI support for NVIDIA GPUs is coming in Hermes Agent, OpenClaw and Perplexity Portable Computer.

    Read the full article at blogs.nvidia.com
    Show the full text · 9 min read

    Frontier intelligence is going local. At IFA 2026, NVIDIA, Microsoft and its partners are teaming up to provide faster inference and new tools that make agents easier to set up and run locally on NVIDIA hardware. New compact NVIDIA RTX Spark Windows PCs are also coming in October to give AI enthusiasts, developers and creators more ways to run capable agents locally and securely. Today’s announcements include: Simplified local AI support for NVIDIA GPUs is coming in Hermes Agent, OpenClaw and Perplexity Portable Computer. Up to 1.9x faster local inference — new llama.cpp and vLLM optimizations are available now directly and through LM Studio and Ollama. NVIDIA PAIR — a Personal AI Router tool that intelligently distributes AI inference across the PCs on a user’s local network. NVIDIA RTX Spark arrives in October — with new Windows PCs from Lenovo and Acer. Electronic Arts, Embark and Ubisoft are among the latest game publishers and developers bringing their blockbuster titles to NVIDIA RTX Spark. Also, August was a busy month for local AI : Nemotron 3.5 Lightning — which can run on NVIDIA RTX PCs, RTX PRO Workstations, DGX Spark and Jetson — is a 30-billion parameter model that has been launched. Get started with Nemotron 3.5 Lightning today. Z.ai’s GLM-5.3-Flash is a multimodal mixture-of-experts (MoE) model that’s bringing agentic AI to DGX Station. Qwen has released Qwen3.8-Flash-Next , an open weight multimodal MoE model, which can run locally on DGX Spark and DGX Station, along with Qwen3.8-27B , a 27-billion-parameter open model optimized for local agentic and coding workloads on NVIDIA GPUs. LTX’s LTX 2.5 is an open-world video generation model optimized for NVIDIA RTX GPUs, DGX Spark and DGX Station, with new NVFP4, FastVideo and ComfyUI enhancements for faster, more memory-efficient local generation. MiniMax-H3 is an open-weight video generation model with synchronized audio that can run locally on NVIDIA GPUs through ComfyUI. FastVideo teamed up with NVIDIA researchers to improve this further by releasing FastH3 — an open-weight, four-step distilled version that improves performance by 7x. Optimized FastVideo recipes for NVIDIA RTX GPUs and DGX Spark are coming soon. Meta’s Muse Glimmer is a 30-billion-parameter open-weight model for coding and agentic workloads that can run locally on GeForce RTX PCs, DGX Spark, DGX Station and Jetson. NVIDIA has also released NVFP4 quantization with DGX Spark support for more memory-efficient local deployment. DeepSeek v4 Flash is a 284-billion-parameter MoE model with 13 billion active parameters that can run locally on 2x DGX Spark cluster and DGX Station. A Simpler Start for Local Agents Getting a local agent up and running with local models required some effort — choosing a model, finding a compatible inference server, dialing in quantization settings and keeping everything updated. That friction is disappearing on RTX and DGX systems. Three of the most widely used agent apps will offer simplified local model setup on Windows, each built on llama.cpp and incorporating NVIDIA’s latest inference optimizations. The new setup experiences are designed to reduce manual configuration and make it easier to get local agents up and running. Last month, Perplexity introduced its Portable Computer agent , giving users a simple way to run Perplexity locally on Linux systems like NVIDIA DGX Spark with the models, orchestration and tools packaged into a single app experience. Perplexity Portable Computer is available on NVIDIA RTX GPUs with at least 24GB VRAM running on Linux, with support on Windows coming soon, bringing that same streamlined setup to a broader group of PC users. Users can run complete workflows locally without consuming credits, while selectively escalating parts of a task to one of 15+ frontier models in the cloud when additional research or reasoning is needed. Portable Computer asks for permission before sending content to the cloud, helping users keep sensitive information on their device. Here’s some example use-cases: Engineering: Review open PRs in a connected GitHub repo and sort them into ready, blocked, stale, and needs review, each tagged with the next step. Docs that fell out of sync with the latest merge get caught and fixed, with a PR opened for the changes. Finance: Point the agent at two years of brokerage summaries, consolidated 1099s and tax returns and have it trace the recurring holdings creating the most avoidable fees and tax drag, with every figure cited to the exact file and page — all without a document ever reaching a chatbot. Startups: Ask why activation went flat, and the agent analyzes the funnel export locally to find where new signups drop off between install and first completed task, then posts the top insights straight to the team’s Slack channel. Try Portable Computer today. Hermes Agent — developed by Nous Research — is a general-purpose agent used by millions that excels at reliability and self-improvement. Model- and provider-agnostic, Hermes is built to run all day on local systems, making RTX PCs, RTX PRO workstations and DGX Spark a natural fit. Configuring a local model in Hermes will provide users with one-click setup across RTX and DGX systems on Windows. The agent will automatically detect the NVIDIA GPU, select an appropriate model and configuration, and run it through integrated llama.cpp with NVIDIA inference optimizations already in place, eliminating manual model downloads and tuning. Support for Linux is coming soon. Once it is running, Hermes works the way it does anywhere else. It uses tools, maintains context across tasks, remembers information between sessions and creates reusable skills over time, allowing the agent to become more capable with continued use. Running the model locally on a GPU keeps performance fast while keeping data on the system. One-click local model setup is available now on Windows, with support coming soon to Linux. Learn more about Hermes Agent . OpenClaw has become one of the defining projects of the open-agent movement — the largest AI project on GitHub, with more than 380K stars and a fast-growing community that’s building tools and skills across research, engineering, project management and everyday productivity. NVIDIA, Microsoft and OpenClaw have been working together to make that experience easier to set up on Windows PCs. To reduce onboarding friction, the OpenClaw Windows App simplifies the process of setting up an optimized local model on any RTX GPU with at least 24GB of VRAM. Learn more in the OpenClaw blog . Faster Inference Gives Local Agents a Boost Inference performance is critical to keeping local agents responsive. NVIDIA is continuing to collaborate with the open-source llama.cpp and vLLM communities to accelerate agentic workloads across local NVIDIA platforms. llama.cpp delivers up to 1.9x higher throughput through kernel optimizations on a GeForce RTX 5090, enhanced speculative decoding techniques and faster prefill. ​ vLLM delivers 1.2x on RTX PRO 6000 Blackwell Workstation Edition and up to 1.4x on two DGX Spark clusters. New XQA attention kernels in FlashInfer and backend optimizations help to accelerate inference across both platforms. These gains are available on the llama.cpp and vLLM inferencing backends. Users can also experience these via the LM Studio and Ollama applications. Tap Idle PCs for More Local AI Compute With NVIDIA PAIR More than half of U.S. households have two or more PCs, and much of that computing power sits idle throughout the day. NVIDIA Personal AI Router (PAIR) is a free, open source software tool that puts those systems to work together for local AI. Agentic workflows often break complex tasks into smaller jobs that can run in parallel, but performance can slow when every request is competing for the same GPU. PAIR automatically discovers compatible PCs on a local network and routes independent inference requests to whichever system has capacity. It works

    Don't miss tomorrow's

    The Daily Pulse in your inbox each morning — sourced and linked.

    How often
    Keep going — across the app