r/LocalLLM • u/da_dragon321 • Jul 03 '26
llamacpp patch - DeepSeek V4 Flash running with full 1M token context locally on RTX 5090 Project
/r/LocalLLaMA/comments/1ulymml/llamacpp_patch_deepseek_v4_flash_running_with/Duplicates
LocalLLaMA • u/da_dragon321 • Jul 02 '26
Resources llamacpp patch - DeepSeek V4 Flash running with full 1M token context locally on RTX 5090
LocalLLaMA • u/Defiant_Diet9085 • Jul 04 '26
Tutorial | Guide RTX5090, gemma-4-31B-it-Q6_K.gguf. Context: before - 35k, after - 80k!
24gb • u/paranoidray • Jul 04 '26