Open-Source AI

Nunchaku 4-Bit Diffusion Comes Natively to Diffusers

Nunchaku 4-Bit Diffusion Comes Natively to Diffusers

Nunchaku Lite in Diffusers announcement graphic showing 4-bit diffusion inference

The most effective 4-bit quantization scheme in image generation just lost its biggest adoption barrier. Hugging Face has integrated Nunchaku-style checkpoints directly into Diffusers, the library announced in a July 23 blog post by Pham Hong Vinh and Sayak Paul — meaning a quantized diffusion transformer now loads with a plain from_pretrained() call, no separate inference engine and no local CUDA compilation required.

That last part matters more than it sounds. SVDQuant checkpoints have delivered the best speed-for-quality trade in low-bit diffusion since the SVDQuant paper introduced the method, but actually running them meant installing the standalone Nunchaku engine, a separate CUDA project with its own pipeline classes and checkpoint format. Plenty of builders who would happily halve their VRAM bill never bothered. The new integration path, dubbed Nunchaku Lite, removes that friction: precompiled kernels are pulled from the Hub via the kernels package the first time you run, and the quantized model keeps the exact module structure of the dense one — so schedulers, LoRA loading, offloading and torch.compile all see a perfectly normal Diffusers model.

Why SVDQuant is different from the quantization you already use

Most quantization backends already wired into Diffusers — bitsandbytes, GGUF, torchao, Quanto — are weight-only: they store weights in low precision and dequantize back up at compute time. That slashes memory but rarely speeds anything up, and can even add latency. SVDQuant instead runs the transformer’s attention and MLP projections with 4-bit weights and 4-bit activations (W4A4). The trick is outlier handling: activation outliers are migrated into the weights, the hardest part of each weight matrix is carried in a small 16-bit low-rank branch, and the residual is quantized to 4 bits. Fused kernels make the low-rank correction essentially free.

Nunchaku Lite exposes two kernel families. The svdq_w4a4 layers (available in INT4 and NVFP4 variants) cover the projections where nearly all compute is spent; awq_w4a16 handles precision-sensitive, memory-bound layers like adaptive-norm modulation. The honest caveat, straight from the announcement: without the original engine’s architecture-specific fused execution paths, Lite cannot fully match standalone Nunchaku — but it still delivers roughly a 30% speedup alongside the same VRAM reduction.

The numbers on real hardware

Hugging Face’s benchmarks on an RTX PRO 6000 (Blackwell) at 1024×1024 put the BF16 baseline at 3.00 s per image with 31.1 GB peak VRAM. The Nunchaku Lite NVFP4 build lands at 2.27 s and 20.6 GB; add torch.compile and the full pipeline drops to 1.68 s — a 1.8× end-to-end speedup. On consumer hardware the story is arguably better: a ready-made checkpoint pairing an NVFP4 transformer with a bitsandbytes NF4 text encoder generates a 1024×1024 image in about 1.7 seconds on an RTX 5090 at ~12 GB peak memory, versus roughly 24 GB for the BF16 pipeline. Quantizing the text encoder alone shaves about 22% off peak VRAM — a reminder that the transformer isn’t the only multi-gigabyte tenant on your GPU.

BF16 versus Nunchaku Lite output comparison from Hugging Face's benchmark
BF16 baseline versus Nunchaku Lite 4-bit output in Hugging Face’s image-quality comparison. Image: Hugging Face

One hard constraint to plan around: NVFP4 checkpoints require an NVIDIA Blackwell GPU — RTX 50 series, RTX PRO 6000 or B200. Older Turing, Ampere and Ada cards (RTX 30/40 series, A100, L40S) use the INT4 variants instead, while Volta and Hopper currently get no 4-bit kernel support at all; the loader validates CUDA capability up front and raises a clear error rather than silently producing garbage.

You can quantize your own models now

The piece that makes this an ecosystem move rather than a convenience feature is the companion diffuse-compressor toolkit. It walks through inspecting which modules will be quantized, running the quantization, packaging the result as a regular Diffusers pipeline, and pushing it to the Hub — meaning new architectures no longer wait on model-specific integration work in the upstream engine. A Nunchaku Lite repository is just an ordinary Diffusers repo with a quantization_config block in the transformer’s config.json; anyone can publish one.

For the local-generation crowd this closes a long-standing gap between what research kernels could do and what the mainstream tooling would load. The 20–30 GB VRAM ask of modern BF16 diffusion transformers put them out of reach of nearly every consumer GPU. With W4A4 checkpoints now first-class citizens in the most widely used diffusion library, the practical floor just dropped to hardware people actually own.

We may earn commission from affiliate links at no extra cost to you. Last updated: Jul 24, 2026.
Jinultimate

Editor of ZBrandCo and the person accountable for what we publish — setting our sourcing standards, fact-checking claims against primary sources, and issuing corrections promptly across AI, open source, and gaming. Reach the desk at editorial@zbrandco.com.