AI

NVIDIA spotlights open models for on-device AI agents

NVIDIA spotlights open models for on-device AI agents

NVIDIA and Local AI Community Fuel Open Source Models and Intelligent Agents | NVIDIA Blog

NVIDIA opened a special-edition “Local AI” blog series on August 11, using the month to spotlight the open models, applications, and tools that let developers run increasingly capable AI agents on their own hardware instead of in the cloud NVIDIA blog.

The hardware story is a burst of new open-weight models tuned for on-device and small-system inference. NVIDIA says Cosmos 3 Edge, a 4-billion-parameter open world model for robotics, autonomous vehicles, and vision AI, is one-quarter the size of Cosmos 3 Nano and runs on device on DGX Spark and Jetson NVIDIA blog. Poolside AI’s Laguna S 2.1, a 118-billion-parameter agentic coding model, ships an NVFP4 checkpoint that NVIDIA says runs locally on a single DGX Spark, while DeepSeek-V4-Flash, a 284-billion-parameter mixture-of-experts model with 13 billion active parameters and a 1-million-token context window, can run on a DGX Station via community GGUF builds NVIDIA blog.

Meta’s Muse Glimmer anchors the coding-and-agent angle. NVIDIA reports the 30-billion-parameter dense open-weight model, built for local agentic AI with a 120K-plus context window, delivers more than 200 tokens per second on an RTX 5090, letting always-on agents process private files and credentials on a single consumer GPU NVIDIA blog.

Thinking Machines Lab’s Inkling-Small is the one model in the roundup not built by NVIDIA. The lab says its 276-billion-parameter mixture-of-experts model activates just 12 billion parameters per token, trained on GB300 NVL72 systems, and runs on a single DGX Station or two DGX Spark systems, with an NVFP4 checkpoint on Hugging Face NVIDIA blog Thinking Machines Lab.

NVIDIA also expanded its own model family with Nemotron 3.5 Lightning, a 30-billion-parameter MoE model that the company says delivers up to 4x faster token generation and 30% faster time to completion than open models in its class, and ships with a routing library, NeMo Switchyard, that NVIDIA says cut benchmark completion cost to roughly one-third of Opus 4.8 alone NVIDIA blog.

Tooling follows the models. NVIDIA Sync’s new Cluster Assistant automatically links two or more DGX Spark systems into a high-speed cluster over ConnectX-7 ports, and DGX Spark will add a native ARM64 Chrome build and a resource monitor later in August. Unsloth Desktop, launching as a fully open-source app for local inference, fine-tuning, and agent workflows, is billed by its maker as the first desktop app that both trains and runs AI models locally NVIDIA blog.

For developers, the result is a widening choice of open models that no longer require a data center: a single RTX 5090 or a small DGX Spark cluster can now host agents that used to need cloud GPUs, with the privacy and cost control that local inference brings. zBrandco’s coverage of NVIDIA’s open-source model router for AI agents examines the company’s separate open routing library, NeMo Switchyard, which aims to mix these models per task instead of renting one large frontier model.

Editorially independent: we accept no payment for coverage and currently use no affiliate links. Read our Editorial Standards and Corrections Policy. Published: Aug 12, 2026.
Jinultimate

Editor of ZBrandCo and the person accountable for what we publish — setting our sourcing standards, fact-checking claims against primary sources, and issuing corrections promptly across AI, open source, and gaming. Reach the desk at editorial@zbrandco.com.