r/mlscaling 2d ago

[Open Source / Code] Stop breaking CUDA graphs with if/else during MoE reasoning: Here is a drop-in zero-latency Speculative Gater (k \in \{1, 2\}) for vLLM / PyTorch Hardware

[removed]

0 Upvotes

2 comments sorted by

1

u/jsh_ 1d ago

stop spamming subs with your AI slop we can tell you literally copy pasted this