Links indicate relevance, not agreement. How to use this site →
DiffusionGemma is an open-weight language model that uses discrete diffusion to generate text in parallel blocks of 256 tokens, achieving around 1,500 output tokens per second on a single GPU—substantially faster than autoregressive models. The model is created by fine-tuning Gemma 4 with a two-stage training pipeline combining supervised fine-tuning and reinforcement learning, while retaining support for thinking mode, multimodal inputs, and long contexts.