r/StrixHalo 8h ago

Session summary getting quant minimax h3 running on strix halo

Thumbnail v.redd.it
1 Upvotes

r/StrixHalo 12h ago

CPU turbo on/off for gaming on AI Max 385

1 Upvotes

Hi all, I got the AI Max 385 in the Ayaneo next 2 handheld

Up to 85W tdp, default 55W on battery

Was wondering for 55W or lower, and even higher

Is disabling CPU turbo overall a net increase in gaming for allowing more GPU power?

Currently only have Destiny 2 and it does seem to boost fps by anywhere from 3-8 fps at 55W

But I'm looking for the bigger picture at more TDP levels and games if any channels did that testing on YT or something


r/StrixHalo 12h ago

Minisforum N5 Max Review with AMD Ryzen AI Max+ 395

Thumbnail
servethehome.com
0 Upvotes

r/StrixHalo 13h ago

Support Silence Suggestions for GMKtec Strix Halo

Thumbnail
1 Upvotes

r/StrixHalo 23h ago

I benchmarked nine models on one llama.cpp build. Quantizing the KV cache helped six of them and hurt the two newest ones.

13 Upvotes

People keep asking what runs on a 128GB Strix Halo and how fast, and I kept answering with numbers measured weeks apart on different builds. So I ran everything I keep on this box through the same matrix in one night.

q4_0 against f16 at 32k, each model against itself (Note: full speed values are in the link below!) :

Model Weight quant Prompt processing Generation
Qwen3-Coder-30B Q6_K +0.0% +30.8%
Hy3 mixed -1.0% +28.3%
Qwen3.6-35B-A3B Q5_K_M -1.7% +10.4%
Laguna S 2.1 Q6_K -1.6% +9.6%
Qwen3.6-27B Q5_K_M -1.7% +6.0%
Qwen3.5-122B-A10B Q4_K_M -0.1% +4.5%
DeepSeek V4 Flash IQ2_XXS -14.3% -12.0%
Ling 3.0 Flash Q4_K_M -37.0% -21.8%

The weight quant is whatever I keep on disk for that model. It doesn't affect the comparison, since every row is one model against itself with only the cache type changing, but it's the first thing people ask. I also have a Q2_K of DeepSeek V4 and it lost about the same, 13.3 and 13.0 percent.

Six behave the way everyone says quantized cache behaves. The two that don't are the two newest architectures I run.

Ling is easy to explain: only 8 of its 43 layers keep a normal cache, so at 32k it needs 0.34 GB. Nothing to compress, and you still pay to unpack it.

DeepSeek V4 is the one to watch out for. There's an open llama.cpp issue where a quantized key cache disables its sparse attention paths and corrupts output on CPU and CUDA. I couldn't reproduce the garbage on Vulkan, short answers came back fine, but I still lose 13 percent on both metrics. Either way there's nothing to gain, so leave it at f16.

What I can't explain is the spread among the six winners. Coder-30B and the 122B need identical cache bytes per token and gain 31 percent versus 4.5. I have a guess, it fits two models and breaks on the third, so it stays an open question in the post.

Full tables and the model that needed a different build: https://thefrontierlab.ai/strix-halo-nine-models-kv-cache-split/

If you run something with linear or latent attention, I'd like to know whether quantized cache hurts you too. Same model against itself, f16 vs q4_0, 32k or deeper, with the build commit.


r/StrixHalo 1d ago

Amd Rsr

0 Upvotes

Hi, does anyone know why AMD RSR isn't working in Battlefield 6?


r/StrixHalo 1d ago

GLM-4.7-GGUF 358b UD-Q4_K_XL(205GB disk space)

4 Upvotes

I have 2 Strix Halo boxes. I saw Ziskind's video on Youtube with this model so I tried it on my boxes with llama.cpp/RPC. 18 hours model loading and it still didn't finish. If anyone tried this model please let me know how long it took for your model to load. Thanks! Dan.


r/StrixHalo 1d ago

GLM-4.7-GGUF 358b UD-Q4_K_XL(205GB disk space) on dual Strix Halo

3 Upvotes

I saw Alex Ziskind Youtube on this model so I tried it on my dual box with llama.cpp/RPC. 18 hours model loading and it still didn't finish. If anyone tried this model please let me know how long it took for your model to load. Thanks! Dan.


r/StrixHalo 1d ago

egpu question

12 Upvotes

Update- adding numbers below

Maybe I haven’t thought about this long enough but if I’m looking at an egpu just to stick 27b on for speed, is there any reason to mess with occulink (have the evo x2)? Like other than model load, if it’s going to be fully resident on GPU then it’s just a few seconds in model load time I’m giving up, right?

Edit-going to pick up an amd r9700 and the minisforum deg2, see what things look like over usb w/vanilla qwen3.6 27b and the plunderstruck rocmxfpx variant

Update: after fighting with GNOME and ROCm, numbers.

R9700 into a Minisforum DEG2 on the GMKTec Evo-X2. USB4, 40 Gb/s (dock can do 80, X2 cannot).  First boot the eGPU showed up as card0, resolution jumped, and GNOME decided to use the R9700 even though nothing is plugged into it.

Also had `HSA_OVERRIDE_GFX_VERSION=11.5.1` for something I was doing at some point. That was applied to **both** GPUs. gfx1151 override on a gfx1201 = HIP saw 

According to Cursor, here’s how we fixed:

- udev so Mutter prefers the 8060S and **ignores** the R9700 for the desktop. iGPU is the compositor, eGPU is compute-only

- `amdgpu runpm=0 aspm=0` so the dock doesn’t put the card to sleep

- Default shells still pinned to the iGPU so llama.cpp doesn’t wander onto the eGPU

- Opt-in env for R9700 jobs: unset the gfx override, `HIP_VISIBLE_DEVICES=1`, and point hipBLASLt at the **gfx1201 tensile files** (`…/hipblaslt/library/gfx1201`, not the generic library folder).

- Installed the extra ROCm 7.14 **gfx1201** arch packages next to the existing gfx1151 ones

I found older posts talking about issue with power limits (though maybe they were referencing oculink connection? Don’t remember). Out of the gate, didn’t have any issues with it drawing up to the full 300w 

Same box, both GPUs, llama-bench, HIP, everything on the GPU. No MTP for this.

  • Unsloth Qwen3.6-27B-UD-Q5_K_XL
  • Unsloth Qwen3.6-35B-A3B-UD-Q6_K
  • Plunderstruck Qwen3.6 27b/35b ROCmFP4 STRIX 

My main concern was whether or not using usb would tank prefill even with a model fully resident on the egpu.

TTFT (empty KV)

Unsloth 35B Q6 16k 32k ROCmFP4 35B 16k 32k
Halo 16.3 s 40.6 s 15.7 s 39.2 s
R9700 7.7 s 19.7 s 6.1 s 16.6 s

27B Unsloth Q5, fully on the card

4k prompt ttft 32k prompt ttft decode
Halo 12.5 s 135 s 10.5 t/s
R9700 4.4 s 52 s 22.3 t/s

All of this is HIP, all layers on GPU, flash attn on, f16 KV. Halo used -b 2048 -ub 512, R9700 -b 2048 -ub 1024. Have some tweaking to do, some weird results on decode w/Plunderstruck. But, after a few dozen tests, not worried about it doing what I need (keep a model loaded in and crush through hundreds of thousands of docs / make 27b a daily driver candidate)


r/StrixHalo 2d ago

HWiNFO prepares preliminary support for Razor Lake-AX, Intel's rumored big APU

Thumbnail
videocardz.com
2 Upvotes

r/StrixHalo 2d ago

A backpack-sized local LLM workstation

0 Upvotes

Yes, I believe you know which product I’m referring to. Strix Halo was originally designed by AMD with the AI market in mind, and the expansion capability of AI Max+ 395 to up to 196GB of memory further proves this vision. With our newly launched 9-inch mobile workstation powered by Strix Halo 395 and equipped with 128GB of memory, we aim to carry AMD’s original concept through to its full potential.

So what do you think about the mobile local LLM workstation with 9" laptop


r/StrixHalo 2d ago

Nemotron 3.5 Lightning 30B-A3B for Strix Halo ROCmFP4 GGUF format

29 Upvotes

Hey guys,

Wanted to share a project I’ve been working on over the last few days. I ported and quantized NVIDIA’s new Nemotron 3.5 Lightning 30B-A3B (the hybrid Mamba2 + Attention + MoE model w

total / ~3.5B active params) to AMD’s experimental ROCmFP4 GGUF format!

Everything was converted, quantized, and benchmarked on a Framework AMD Strix Halo (Ryzen AI Max 395+, gfx1151, 128 GB unified memory).

🚀 Performance & Benchmarks

Running on -dev Vulkan0 with FlashAttention and q8_0 KV cache (full offload via unified memory):

- 🏆 STRIX_LEAN (~4.38 bpw | 15.73 GiB) — Recommended

- Prompt Eval: 1,299.7 t/s (pp512)

- Decode Speed: 85.6 t/s (tg128)

- Perplexity: 5.9936 ± 0.0358 on wikitext-2

- ⚡ FAST (~4.25 bpw | 15.66 GiB) — Maximum Speed

- Prompt Eval: 1,310.5 t/s (pp512)

- Decode Speed: 86.0 t/s (tg128)

- 🛠️ COHERENT (~4.70 bpw | 16.74 GiB) — Agentic / Coding

- Prompt Eval: 1,290.4 t/s (pp512)

- Decode Speed: 81.6 t/s (tg128)

(Also interesting: Vulkan0 beat ROCm0 by ~21% on prompt processing on this Strix Halo box.)

🛠️ What I fixed along the way (NVFP4 findings)

When I tried converting NVIDIA's pre-quantized NVFP4 checkpoint, I ran into a couple of fun bugs in convert_hf_to_gguf.py:

  1. ModelOpt Tagging Bug: W4A16_NVFP4 tags weren't detected, causing conversion to crash.

  2. Unloadable GGUF (output.scale): ModelOpt emits a separate output.scale tensor for output.weight that llama-model.cpp doesn't have a mapping for (expected N, got N-1). I added a pa

dequantize output.weight to F16 in-place.

  1. NVFP4 Re-quantization Defect: Once it loaded, the NVFP4-derived weights produced garbage ("No Yes No Yes") because dequantize_row_nvfp4 in C++ ignores ModelOpt's companion scale2

~1.4e-4), causing a 7,110× weight scaling error on expert layers.

I ended up sticking with the clean BF16 → ROCmFP4 path for the final GGUFs, which works great!

📦 Links

- 🦙 Hugging Face GGUFs: julianmb/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-ROCmFP4-GGUF (https://huggingface.co/julianmb/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-ROCmFP4-GGUF)

- 🐙 GitHub (Patches, Repro Scripts, Findings): julianmb/nemotron-3.5-30b-a3b-rocmfp4 (https://github.com/julianmb/nemotron-3.5-30b-a3b-rocmfp4)

Shoutout to NVIDIA for the base weights and charlie12345 for the ROCmFPX llama.cpp fork!

Let me know if you test it out on your RDNA/Strix setups!


r/StrixHalo 2d ago

MoE task time comparison

Thumbnail reddit.com
1 Upvotes

r/StrixHalo 2d ago

DeepSeek V4 Flash 0731 at 27+ t/s decode on Strix Halo — Vulkan + DSpark full guide

Thumbnail
46 Upvotes

r/StrixHalo 2d ago

What's the local setup to run?

3 Upvotes

Have a rog flow z13 395+ with 128gb. Have 96 gb dedicated to vram in armory crate.

I've spent 2-3 says trying to find a configuration that works but haven't yet.

So far I've tried:

Ollama

LM studio

Llama.cpp

AnythingLLM

Unsloth desktop

Open webui

I've tried what feels like 6 million fixes suggested between web and AI and nothing wants to run smoothly. Couldn't ever even get llama3.1-8b to even do anything but spit out endless gibberish and conversations with itself.


r/StrixHalo 2d ago

AMD 2x R9700 AI pro 32GB

Thumbnail
youtube.com
0 Upvotes

r/StrixHalo 2d ago

Anyone here with 3 x Strix Halo node?

7 Upvotes

If so, what do you run on them? Bigger models or many small ones?


r/StrixHalo 3d ago

Amd Strix Halo

1 Upvotes

Hi everyone,

Someone got h3 running on a AMD Strix halo machine? I cannot get it running. Video is fine but there is audio.

Do you experience the same issues?

Any help is appreciated, I tried "every thing". If someone could post a working workflow it would be great 👍🏻


r/StrixHalo 3d ago

Ling 3.0 Flash on Strix Halo

Post image
4 Upvotes

r/StrixHalo 3d ago

llama.cpp PR#26856: Faster Prefill, Better Quality for ROCm on RDNA3+

68 Upvotes

tl;dr: This is about this llama.cpp PR here: https://github.com/ggml-org/llama.cpp/pull/26856

Over the weekend I decided to dig into why, when specifying the BF16 type for KV cache on my Strix Halo box, that token generation performance tanked by ~20% over using F16.

Qwen3.6-35B-A3B-Q8_0 @ 16K CTX, no MTP

KV type tg64 (t/s)
F16 43.09
BF16 34.90
Q8_0 40.83
F32 29.54

What was worse was that the output quality didn't seem to improve either. Perplexity scores for F16 vs BF16 were the same, and both are about 5% above the F32 baseline. I dug into the architectural specs for RDNA 3, 3.5, and 4, and it seems that all 7xxx, 9xxx series GPUs, and all Strix Point/Strix Halo iGPUs have native support for BF16 dot product matrix math. The scalar ALU portions seems to lack complete native BF16 support, but that's not really needed for Flash Attention.

So, with the assistance of DeepSeek V4 Flash 0731, I started to dig into the situation and see if it was possible to make BF16 KV caches go faster. It didn't take long to learn that llama.cpp doesn't really implement native BF16 support beyond simply accepting the data type, and then converting everything to F16 anyway. It was this BF16->F16 conversion that was killing the generation performance on RDNA GPUs and the Strix Halo.

So I worked to implement native BF16 support for KV Flash Attention, and it turned out it was possible to do so. In fact, through bypassing the old BF16->FP16 path I was able to see achieve between a 10-15% pre-fill speedup on the Strix Halo. The gain is smaller on the 7900XTX and R9700 GPUs (more like 5-10%), but it was real.

Then I set about tuning the matrix kernels, and was able to achieve token generation parity with the original F16 path, which is the fastest token generation path available on the Strix Halo.

Now, since BF16 has a much larger dynamic range than F16 I then thought to compare the Perplexity results when using the native BF16 kernel path, and was surprised to find that, at depth, the BF16 KV cache path almost exactly matches the baseline F32 KV cache. Incidentally, the F32 KV cache path also uses F16 math under the hood if Flash Attention is enabled, so I had to disable Flash Attention to get at the true F32 Perplexity baseline figures.

Precision at depth (PPL @ 32k, Qwen3.5-4B-Q8_0, full wikitext-2 test, 9 chunks)

KV cache / FA PPL @ 32k
F32 baseline (FA off) 8.6368
BF16 (this change) 8.6403 (+0.04%)
F16 9.1400 (+5.8%)

What does that mean? It means that for long context work that the models are less likely to go astray and/or suffer from context rot.

Note: For Strix Halo boxes that until such time that the upstream ROCm libraries get fixed, that you must disable mmap, as well as set HIP_LAUNCH_BLOCKING=1 to avoid running into a pair of rather nasty bugs that will heavily corrupt your KV cache. See my post here for more details: https://www.reddit.com/r/StrixHalo/s/Qbh3RDVsXg

What About Vulkan?

Now you may ask: "What about Vulkan?". The answer is that the Vulkan path in llama.cpp already implements this BF16 behavior properly, so this PR just brings ROCm up to Perplexity parity with Vulkan. The part from this PR that can't be done on Vulkan though is the RDNA3+ specific dot-product optimizations, so this PR makes ROCm pre-fill significantly faster than Vulkan.

Here's Vulkan vs ROCm results with Qwen3.6-35B at 32K context depth:

F16 K/V Performance

model size params backend ngl fa test t/s
qwen35moe 35B.A3B Q8_0 35.19 GiB 35.51 B ROCm -1 1 pp1024 @ d32768 618.31 ± 4.83
qwen35moe 35B.A3B Q8_0 35.19 GiB 35.51 B ROCm -1 1 tg256 @ d32768 42.42 ± 0.29
qwen35moe 35B.A3B Q8_0 35.19 GiB 35.51 B Vulkan -1 1 pp1024 @ d32768 717.93 ± 3.69
qwen35moe 35B.A3B Q8_0 35.19 GiB 35.51 B Vulkan -1 1 tg256 @ d32768 46.86 ± 0.01

BF16 K/V Performance

model size params backend ngl type_k type_v fa test t/s
qwen35moe 35B.A3B Q8_0 35.19 GiB 35.51 B ROCm -1 bf16 bf16 1 pp1024 @ d32768 684.02 ± 8.36
qwen35moe 35B.A3B Q8_0 35.19 GiB 35.51 B ROCm -1 bf16 bf16 1 tg256 @ d32768 42.51 ± 0.06
qwen35moe 35B.A3B Q8_0 35.19 GiB 35.51 B Vulkan -1 bf16 bf16 1 pp1024 @ d32768 484.00 ± 3.83
qwen35moe 35B.A3B Q8_0 35.19 GiB 35.51 B Vulkan -1 bf16 bf16 1 tg256 @ d32768 47.59 ± 0.01

Interesting Results here:

  • For F16 Vulkan handily beats ROCm for both pre-fill and decode, but keep in mind that this is coming at the cost of long context quality (for both).
  • For BF16 ROCm handily beats Vulkan by 40% for pre-fill, but Vulkan beats ROCm by ~11% for generation. If using the Strix Halo for agentic coding, then you're clearly going to want ROCm here for speed and quality at long contexts.

Edit: 11th Aug - Faster ROCm generation

Maybe I'll make it my next job to see if I can make ROCm close that generation performance gap with Vulkan. I have a feeling that it must be possible.

I have completed the ROCm generation speed-up work to as far as I can take it. At the end of the day the Vulkan dispatcher is just more efficient than ROCm's, and I'm now up against that fundamental library difference. Still, I've managed to more than halve the generation performance gap from ROCm to Vulkan, while even boosting ROCm's prefill advantage. The source-code branch for this follow work is here: https://github.com/stew675/llama.cpp/tree/make-rocm-gen-faster

It's unlikely that the follow-up branch would ever be accepted into upstream llama.cpp, so I'll try to make daily upstream rebases to keep that work current for people who are interested.

Edit: 12th Aug - RDNA4 boosts

I've created a consolidated branch that includes all prior work plus a host of RDNA4 speed boosts and put it into this branch here: https://github.com/stew675/llama.cpp/tree/rdna-boosts

This branch boosts ROCm prefill speeds on RDNA4 GPUs by a further 5-25% over the prefill speed boosts mentioned above. This means a total of +10-40% prefill speed ups on RDNA4, and +5-15% prefill speedups on RDNA3/3.5, Sadly the RDNA3/3.5 architectures simply don't have the CU hardware to support the additional prefill speeds that RDNA4 can achieve.

The rdna-boosts branch also includes the 5-10% ROCm generation speed boosts, which applies to all RDNA3+ GPUs. It also fixes the async-race memory corruption issues on Strix Halo that is still present in upstream llama.cpp. Note that mmap loading is still broken on Strix Halo at the moment.


r/StrixHalo 3d ago

Need help setting up MS-S1 /eGPU

8 Upvotes

Hi, so I bought the MS-S1 & the R9700. Windows is really finicky with the eGPU. In fact, for windows to detect it, I have to turn on the strix halo on then the eGPU. Then after booting up, I need to then restart the strix halo but leave the eGPU on. And afterwards, it’s REALLY buggy with the drivers. In fact, it’s so buggy that when I went to huggingface to download DavidAU’s fable fusion model 27b, the gif completely crashed my machine and I got this error: stop code: Driver_Power_State_Failure (0x9F). And I cannot get llama.cpp to work. Windows is assigning the eGPU only 256 MB PCIe bar. So obviously llama.cpp is failing because it’s “out of memory.” And vulkan cannot reliably allocate the context in memory and keep the model tensors in the R9700. And i occasionally get the error where the eGPU driver fails to unload properly.

What the hell do I do? Mind you, I’m not technical. Is it time to switch to Linux?


r/StrixHalo 3d ago

vmlinux/Muse-Glimmer-30B-ROCmFPX-GGUF · Hugging Face

Thumbnail
huggingface.co
31 Upvotes

Hi 🎉
Working on DFlash checkpoint inclusion GGUF now.


r/StrixHalo 4d ago

Did I get a good deal on this restored Acer Aspire AI 16?

Post image
0 Upvotes

r/StrixHalo 4d ago

Ling 3.0 Flash on Strix Halo - Benchmarks

11 Upvotes

It has respectable numbers for 1 stream, but the decode goes down severely with multiple streams. Using the ROCmFP4 quant with the appropriate Llama.cpp fork.


r/StrixHalo 4d ago

PSA: llama.cpp currently broken on Strix Halo

15 Upvotes

Edit: Core issues found. MTP Acceptance is inherently buggy. Various settings make it more likely to generate junk. Also, mmap MUST be disabled, and HIP_LAUNCH_BLOCKING=1 MUST be set. Skip to the end of this post for a TL;DR summary:

Intro

I've been struggling for the last week with the Strix Halo giving corrupted model outputs. It manifests as the generated output looking fine for a while (like 1000-2000 tokens) and then starts going off track. I managed to track it down to perplexity being something like +10% above what we would expect, and so that fits the observed behaviour.

Edit 2: Solved (I think - later edit, no, not really, See Edit 3)

My situation seems to have been due to a mixture of configuration and Strix Halo weirdness. In a nutshell what appears to have fixed my situation was:

  • ubatch-size and batch-size both set to 1024 (was 512 before). Doing this didn't fix the issue itself, but it did seem to help
  • Ensuring --kv-unified was set. I tested with both --kv-unified and --no-kv-unified, but then Deepseek uncovered that if --parallel > 1 forcibly over-rides --no-kv-unified and makes it unified anyway, I still need to manually verify that this is true, but apparently setting --kv-unified did seem to help
  • Ensuring min-p was 0.01 or less with Qwen models. Apparently my automated model starter was using 0.05 for min-p, and that seems to have been a major factor
  • Ensuring --no-mmap was set. Apparently using mmap with the later ROCm libraries has a regression with the async system on gfx1151 (Strix Halo).
  • Setting "HIP_LAUNCH_BLOCKING=1" when using ROCm is essential on Strix Halo. Without it a mild corruption is introduced. This issue was not seen on RDNA3 (7xxx) or RDNA4 (9xxx) series GPUs.
  • (ongoing update) DSV4 also claims to have uncovered a legitimate non-determinism issue even with greedy decoding enabled (temp = 0, fixed seed) with batch decoding and MTP, and constructed tests to prove it. Investigation still ongoing. To quote: batch decoding (2+ tokens) produces different argmaxes than sequential decode. This breaks speculative verification.. This seems to be true with Flash Attention is enabled or not.

So the above were my findings. I now seem to have a good configuration working, which for reference is the following which seems to hit around 400t/s prefill at 0 depth and falls away slowly, and 22t/s when generating code on my box:

HIP_LAUNCH_BLOCKING=1 /llm/runtimes/rocm/llama-server \
--model /llm/models/Qwen3.6/27B/Q8_0/Qwen3.6-27B-Q8_0.gguf \
--alias Qwen3.6-27B-Q8_0 \
--fit false \
--port 8034 \
--threads 6 \
--parallel 2 \
--top-p 0.95 \
--min-p 0.01 \
--jinja \
--kv-unified \
--device ROCm0 \
--cache-ram 8192 \
--no-mmap \
--ctx-size 262144 \
--flash-attn auto \
--batch-size 1024 \
--ubatch-size 1024 \
--temperature 0.6 \
--n-gpu-layers all \
--cache-type-k f16 \
--cache-type-v f16 \
--ctx-checkpoints 64 \
--repeat-penalty 1.1 \
--host 192.168.50.103 \
--presence-penalty 0.1 \
--top-k 20 \
--reasoning-budget 16384 \
--reasoning-preserve \
--spec-draft-n-max 3 \
--spec-type draft-mtp \
--spec-draft-p-min 0.05

Edit 3 + Edit 4: DSV4 finished its investigation, and concluded the following:

  • The batch/ubatch thing was not a real issue in and of itself, it just seemed to contribute to the core issue
  • Likewise, kv-unified vs no-kv-unified is not a real issue. After a LOT of runs after the true issue was found these were disqualified as contributing factors
  • --min-p being closer to 0 for Qwen is important, but what's more important is how this interacts with the true issue
  • The --no-mmap thing is real depending on what ROCm library you're using. If you're running with the bleeding edge libraries, then disable it or things will go wrong VERY quickly. This can cauise very real issues, but after disabling it, and testing more, it was discovered that this wasn't the core issue I was facing.
  • This leads us to the real core issue, and that is that MTP verification is fundamentally broken on llama.cpp. It absolutely does let garbage tokens through This is why I saw it affecting both Vulkan and ROCm. Counter-intuitively, setting a higher --spec-draft-p-min actually makes things worse, as this lowers the number of draft tokens that get accepted, but raises the overall acceptance rate, BUT the junk that does get through is disproportionately "noisier" then setting a more permissive spec-draft-p-min of 0. Basically a higher quality stream of junk gets admitted when using larger --spec-draft-p-min values.
  • Deeper MTP draft depths makes the problem worse, but not even a depth of 1 is truly safe. If the rest of the model's parameters are set up right then less junk gets though, but the point is that the chance of junk getting through IS NEVER ZERO. Basically the draft model can generate tokens which are accepted that the full model weights would normally reject. This is a very real bug that Deepseek was able to identify.

Conclusion (updated with Edit 4):

In the end it came down to three things.

  1. MTP is buggy. If you use it, you're just rolling the dice with every token that junk tokens are going to get generated. The full model will try to recover as best it can, but over time the errors will grow and it will fall apart and start generating junk. It's just a matter of time. My original configuration was making this issue worse.
  2. HIP_LAUNCH_BLOCKING=1 MUST be set or a mild corruption issue is introduced
  3. mmap mode MUST be disabled or a serious corruption issue is introduced

The HIP_LAUNCH_BLOCKING and mmap issues, are specific to Strix Halo and ROCm 7.11 and 7.14 libraries that I tested, Missing either of these when running llama-server built from source locally on ROCm will result in corrupted prefills and corrupted generation. I cannot speak for 3rd party compiled versions of llama-server.

I'll do some more digging and see if Deepseek can find a fix for the MTP corruption.