← r/LocalLLaMA
▲
2
+1
2👁
r/LocalLLaMA · u/Impossible_Art9151 · 24h ago

struggling with llama.cpp 2 x dgx spark mtp files start command (unsloth)

Hi all,

having searched google and asked several AIs without success, maybe s.o. can help.
I have 2 x dgx spark in a cluster. deepseek-flash is running successful.
Now I want to test qwen3.8-flash-next from unsloth in the mtp version.

Following start-command runs into a dgx-stall:

./llama-server -hf unsloth/Qwen3.8-Flash-Next-GGUF:Q8\_0 -ngl 999 -ngld 999 --load-mode none --fit off -fa on --host 0.0.0.0 --port 8090 --ctx-size 256000 --parallel 1 --chat-template-kwargs '{"preserve\_thinking": true}' -sm layer --cache-ram 0 --spec-type draft-dspark --spec-draft-n-max 2 --reasoning on --seed 3407 --temp 1.0 --top-p 0.95 --top\_k 20 --min\_p 0.0 --presence\_penalty 0.0 --repeat\_penalty 1.0 --rpc 10.10.188.10:50052

There are two mtp files:
mtp-Qwen3.8-Flash-Next-shared-Q8\_0.gguf
mtp-Qwen3.8-Flash-Next-Q8\_0.gguf

I can't figure out how to use them, start the server properly.
Help appreciated!

2 0 2 10/8 11:29 10/8 19:30 UTC
scorecomments2 sightings
first seen 2026-10-08 11:29 UTClast seen 2026-10-08 19:30 UTCscore then 1score now 2gained +1sightings 2
open on reddit ↗ 💬 11 (+2)