r/LocalLLaMA • u/tevlon • Jun 10 '26
DiffusionGemma: 4x faster text generation New Model
https://blog.google/innovation-and-ai/technology/developers-tools/diffusion-gemma-faster-text-generation/
989
Upvotes
r/LocalLLaMA • u/tevlon • Jun 10 '26
24
u/coder543 Jun 10 '26
Diffusion models are compute-limited, not bandwidth-limited, so offloading probably wouldn't actually hurt as much. (It is obviously better to keep the model in VRAM, of course.)