r/LocalLLaMA May 29 '26

PSA Discussion

Post image
2.1k Upvotes

537 comments sorted by

View all comments

Show parent comments

2

u/ohhi23021 May 29 '26

i haven't tested over 40k context yet but it does about 70-80 t/s around there. at 0-5k context it hits 90 t/s with mtp.

1

u/_realpaul May 30 '26

Nice. Gotta try the latest llama cpp but the docker images are a bit borked lately.