95+ TPS through 100K generated for qwen3.8 27b, 262K ctx, on a single 3090
Hello everyone! A little while back I posted about LlamAmpere, a fork of Llama.cpp with Ampere-specific improvements (though it is caught up to main and will support other hardware, too).
Thank you to everyone that tried it out and shared back their results across the 30xx cards. I'm happy to share I've pushed v0.4 out this morning. On the 4.6bpw model tested, speeds improved \~10% vs the last version while also improving the max context by 10%+ (technically, it can go above 262K, but I have not tested any custom kernels or graphs to support YaRN).
The closest competition comes from vLLM, keeping within <10%, but does so with lower maximum context. It is significantly faster than other llama.cpp options tested.
https://preview.redd.it/u70q7lew1bsh1.png?width=1080&format=png&auto=…
[](https://preview.redd.it/95-tps-through-100k-generated-262k-ctx-on-a-single-30…)
There's also a number of other improvements for other quants/formats, with EXL3 seeing significant speed up (\~80% the speed of the 4-XS-M quant tested). It has a slightly lower KLD, but not a range I have found stat significance for at the task level, so I sticking with the XS-M model for now (built on top of Swift-qwen's distill, which is far more token efficient than the stock train for \~1% performance loss). the 4.3 bpw EXL3 model does provide a bit more room if you are interested in 2+ concurrent predictions. Improvements in this format and the IQ2/3 codebook quants will be most useful for people on 12/16/20 GB setups. These measurements are at temp=1, vs some of the vanity speeds you will see people claim with temp=0 and/or short generations.
As always, please share your results + config details so I can keep improving!
fork is here: https://github.com/JakeATX/llamAmpere/blob/main/QWEN\_AMPERE.md#build-and-run**
model used here:
https://huggingface.co/jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4\_XS-M-GGUF**
Build command:
git clone -b v0.4 https://github.com/JakeATX/llamAmpere.git
llamAmpere
cmake -S . -B build-sm86 -DCMAKE\_BUILD\_TYPE=Release -DGGML\_CUDA=ON -DCMAKE\_CUDA\_ARCHITECTURES=86
cmake --build build-sm86 -j8 --target llama-server
Build + launch (linux):
git clone -b v0.4 https://github.com/JakeATX/llamAmpere.git
llamAmpere
cmake -S . -B build-sm86 -DCMAKE\_BUILD\_TYPE=Release -DGGML\_CUDA=ON -DCMAKE\_CUDA\_ARCHITECTURES=86
cmake --build build-sm86 -j8 --target llama-server
curl -L -o ATX-Swift-Qwen3.8-27B-Uncensored-IQ4\_XS-M.gguf \\
https://huggingface.co/jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4\_XS-M-GGUF/resolve/main/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4\_XS-M.gguf
\-m ATX-Swift-Qwen3.8-27B-Uncensored-IQ4\_XS-M.gguf -c 262144 \\
\-ngl 99 -fa on -ctk turbo5 -ctv turbo4 -b 4096 -ub 1024 -t 8 -tb 8 --parallel 1
cd
./build-sm86/bin/llama-server
Previously, people had expressed concern over quantizing KV cache, and TQ specifically. The TL;DR on that is that any reasonable KV quantization strategy (at least for hybrid attention models like Qwen) is going to be swamped by quantization of the weights. The KV quant we're using here (TQ5/TQ4) is less than 1/3 of the KLD we see when moving from 8 bit weights to 4.6 bit weights (and the KLD is only partially additive, so some of the incremental errors cancel out). There was no statistical significance when testing this KV quant at the task level against 8/8 kv (just trivial variations in sentence length). I will be adding KVaRN in the next release, but with a better codec than currently available elsewhere, so it requires a bit more testing before release.
v0.5 will be focused primarily on the 12GB cards, but this should have generation-wide speed ups, so even if you're not on 24GB, please share your results.
Enjoy!