← r/LocalLLaMA
▲
0
 
8👁
r/LocalLLaMA · u/knighty1981 · 7d ago

2x 3090 in server chassis, upgrade time

I've got 2x 3090 in a supermicro gpu server chassis

Supermicro SYS-4029GP-TRT (will take 8 gpu, but it's only pcie3)
Ubuntu 26.04.1 LTS, 2x Xeon Gold 6230R @ 2.10 GHz, 230gig ram, 2×RTX 3090 24 GB

running huihui-27b 256k context, qwen3.8-27b 64k context, qwen3.8-27b 256k context

using opencode remotely

I've only used the 256k context models, it's fast enough to me

mostly have it doing admin work for me, so it setup a webserver on another server that runs a route planner than it made, had it do a bunch of stuff on my home assistant setup, it's not running yet (waiting on hardware) but I've had it design a voip system using free pbx and whisper to listen in and show prompts on screen (customer details from database it's build etc. tc.)

it's done a load of stuff pulling info from thousands of excel delivery sheets / invoices and summarised them for me / shown trends, bunch of research into competitors (basic summary) etc.

mostly billy basic stuff

over the last week I've had it organise my media server (synology nas) and setup prowlarr/radarr/sonarr/qbittorrent all to run on a vpn (I tried this myself before but got frustrated with it and gave up) - it's been going about 3 days doing this... a lot of slow stuff because it's waiting for the nas to run tasks etc. but it's done a lot of things wrong too, had to go back and change settings, or it's trying to change a setting (over ssh) and using the wrong commands etc. etc. (obv. waiting for input from me too)

running 256k context which it's had to compress a bunch of times

part of this is on me - if I'd known in advance I'd have split it into smaller tasks and had it plan more in advance

as I understand it, running over 256k context is a bad idea because it'll hallucinate more/get stuck in loops?

so... anyone have any hardware upgrade advice? I don't want to spend crazy money, I could get 2 more 3090 so split the model over 4 cards to run faster, or run different models on different cards - I really like the idea of a council of ai but from googling I don't think we're quite there yet?

I could run larger models, does it make that much difference? things are moving so fast when I search for info stuff from 6 months ago is out of date!

I'm not sure if pcie3 will kill performance running more cards with models split over them?

1 0 0 10/3 06:28 10/7 05:39 UTC
scorecomments8 sightings
first seen 2026-10-03 06:28 UTClast seen 2026-10-07 05:39 UTCscore then 0score now 0gained 0sightings 8
open on reddit ↗ 💬 14 (+5)