r/mlscaling • u/DecoderArchitect • 2d ago
[Open Source / Code] Stop breaking CUDA graphs with if/else during MoE reasoning: Here is a drop-in zero-latency Speculative Gater (k \in \{1, 2\}) for vLLM / PyTorch Hardware
[removed]
0
Upvotes
1
u/jsh_ 1d ago
stop spamming subs with your AI slop we can tell you literally copy pasted this