Make Volta Fast Again
For those who have V100 cards, I wanted to point you to 1Cat-vLLM, a vLLM fork that enables optimized serving for these cards. Showing stats for Qwen3.6-35b comparing a Strix Halo with a hughly optimized llama.cpp fork (pwilkin) and the V100 with 1Cat. It’s not apples to apples, but I decided to show the raw numbers from llama-benchy so folks get an idea of the performance. IMO this is still very good for 10 year old GPUs. Welcome any other suggestions for optimization!
scorecomments12 sightings
first seen 2026-10-03 06:32 UTClast seen 2026-10-04 15:47 UTCscore then 30score now 28gained -2sightings 12
open on reddit ↗
💬 82 (-1)