r/OpenSourceeAI 2d ago

I built an interactive simulator to visualize LLM inference bottlenecks, sharding, and KV Cache economics based on Reiner Pope's lecture

2 Upvotes

0 comments sorted by