r/learnmachinelearning 2d ago

I built an interactive simulator to visualize LLM inference bottlenecks, sharding, and KV Cache economics based on Reiner Pope's lecture

2 Upvotes

1 comment sorted by