← r/LocalLLaMA
▲
28
-2
12👁
r/LocalLLaMA · u/Miserable-Dare5090 · 14d ago

Make Volta Fast Again

post image

For those who have V100 cards, I wanted to point you to 1Cat-vLLM, a vLLM fork that enables optimized serving for these cards. Showing stats for Qwen3.6-35b comparing a Strix Halo with a hughly optimized llama.cpp fork (pwilkin) and the V100 with 1Cat. It’s not apples to apples, but I decided to show the raw numbers from llama-benchy so folks get an idea of the performance. IMO this is still very good for 10 year old GPUs. Welcome any other suggestions for optimization!

31 0 28 10/3 06:32 10/4 15:47 UTC
scorecomments12 sightings
first seen 2026-10-03 06:32 UTClast seen 2026-10-04 15:47 UTCscore then 30score now 28gained -2sightings 12
open on reddit ↗ 💬 82 (-1)