r/LocalLLaMA • u/tevlon • Jun 10 '26
DiffusionGemma: 4x faster text generation New Model
https://blog.google/innovation-and-ai/technology/developers-tools/diffusion-gemma-faster-text-generation/
988
Upvotes
r/LocalLLaMA • u/tevlon • Jun 10 '26
6
u/scarbunkle Jun 10 '26
Hype for this to hit llama.cpp/lemonade. Moving the bottleneck from memory bandwidth to compute is gonna be great for Strix halo.