EMNLP 2026
Nemotron-Labs-Diffusion: A Tri-Mode Language Model Unifying Autoregressive, Diffusion, and Self-Speculation Decoding
Nemotron-Labs-Diffusion unifies autoregressive, diffusion, and self-speculation decoding within a single language model trained with a joint AR-diffusion objective. The resulting 3B, 8B, and 14B model family can switch decoding modes across deployment settings, delivering strong efficiency across different serving concurrency levels and outperforming existing multi-token prediction methods. By combining the accuracy of autoregressive generation with diffusion-based parallelism, Nemotron-Labs-Diffusion-8B achieves 4x higher throughput than Qwen3-8B on SPEED-Bench using SGLang on an NVIDIA GB200 GPU.