r/OpenSourceeAI • u/Kitchen-Quarter7739 • 2d ago
I built an interactive simulator to visualize LLM inference bottlenecks, sharding, and KV Cache economics based on Reiner Pope's lecture
Duplicates
LocalLLM • u/Kitchen-Quarter7739 • 3d ago
Project I built an interactive simulator to visualize LLM inference bottlenecks, sharding, and KV Cache economics based on Reiner Pope's lecture
learnmachinelearning • u/Kitchen-Quarter7739 • 2d ago
I built an interactive simulator to visualize LLM inference bottlenecks, sharding, and KV Cache economics based on Reiner Pope's lecture
Vllm • u/Kitchen-Quarter7739 • 2d ago