struggling with llama.cpp 2 x dgx spark mtp files start command (unsloth)
Hi all,
having searched google and asked several AIs without success, maybe s.o. can help.
I have 2 x dgx spark in a cluster. deepseek-flash is running successful.
Now I want to test qwen3.8-flash-next from unsloth in the mtp version.
Following start-command runs into a dgx-stall:
./llama-server -hf unsloth/Qwen3.8-Flash-Next-GGUF:Q8\_0 -ngl 999 -ngld 999 --load-mode none --fit off -fa on --host 0.0.0.0 --port 8090 --ctx-size 256000 --parallel 1 --chat-template-kwargs '{"preserve\_thinking": true}' -sm layer --cache-ram 0 --spec-type draft-dspark --spec-draft-n-max 2 --reasoning on --seed 3407 --temp 1.0 --top-p 0.95 --top\_k 20 --min\_p 0.0 --presence\_penalty 0.0 --repeat\_penalty 1.0 --rpc 10.10.188.10:50052
There are two mtp files:
mtp-Qwen3.8-Flash-Next-shared-Q8\_0.gguf
mtp-Qwen3.8-Flash-Next-Q8\_0.gguf
I can't figure out how to use them, start the server properly.
Help appreciated!