I have an ESC4000 G3 with 8x T4s in it - what is the fastest way I can deploy Qwen3.5-9B for about 10-15 users concurrently: currently using llama.cpp
Hi All:
I have an Asus ESC 4000 G3 with 128 GB DDR4 RAM - I tried putting in my V620s but couldn’t put more than 2. Sadly, pivoted to T4s, these are 72 watt passive cards and I thought I could use them like how I use the Mi50 32GB - but I was very wrong.
It has no support from Nvidia when it comes to latest fp8 emulation, reason being that it lacks resources. I am not sure if special vLLM repos exist, but I am trying to serve Qwen3.5-9B-AWQ-INT8 with vLLM or something faster than llama.cpp.
I currently have 8 separate instances and they’re serving Qwen3.5-9B-Q5\_K\_XL at 87k per slot and there are 44 slots.
So, my agentic (non-coding) harness works, its able to fetch emails, summarize documents, and do a lot for me and my team, but the concern is latency. Each slot continuously generates 25-40 tps (depending on the question), and essentially prefill is at 1400 tps, so I am not sure what my bottleneck is.
VLLM on my Mi50 32GB and AWQ-INT8 (cyankiwi’s model) is very fast.
I am wondering, does anyone have any pointers? I am really not in the position to spend money, exhausted everything for the next 6 months to a year already.
I would greatly appreciate your help.