2x 3090 in server chassis, upgrade time
I've got 2x 3090 in a supermicro gpu server chassis
Supermicro SYS-4029GP-TRT (will take 8 gpu, but it's only pcie3)
Ubuntu 26.04.1 LTS, 2x Xeon Gold 6230R @ 2.10 GHz, 230gig ram, 2×RTX 3090 24 GB
running huihui-27b 256k context, qwen3.8-27b 64k context, qwen3.8-27b 256k context
using opencode remotely
I've only used the 256k context models, it's fast enough to me
mostly have it doing admin work for me, so it setup a webserver on another server that runs a route planner than it made, had it do a bunch of stuff on my home assistant setup, it's not running yet (waiting on hardware) but I've had it design a voip system using free pbx and whisper to listen in and show prompts on screen (customer details from database it's build etc. tc.)
it's done a load of stuff pulling info from thousands of excel delivery sheets / invoices and summarised them for me / shown trends, bunch of research into competitors (basic summary) etc.
mostly billy basic stuff
over the last week I've had it organise my media server (synology nas) and setup prowlarr/radarr/sonarr/qbittorrent all to run on a vpn (I tried this myself before but got frustrated with it and gave up) - it's been going about 3 days doing this... a lot of slow stuff because it's waiting for the nas to run tasks etc. but it's done a lot of things wrong too, had to go back and change settings, or it's trying to change a setting (over ssh) and using the wrong commands etc. etc. (obv. waiting for input from me too)
running 256k context which it's had to compress a bunch of times
part of this is on me - if I'd known in advance I'd have split it into smaller tasks and had it plan more in advance
as I understand it, running over 256k context is a bad idea because it'll hallucinate more/get stuck in loops?
so... anyone have any hardware upgrade advice? I don't want to spend crazy money, I could get 2 more 3090 so split the model over 4 cards to run faster, or run different models on different cards - I really like the idea of a council of ai but from googling I don't think we're quite there yet?
I could run larger models, does it make that much difference? things are moving so fast when I search for info stuff from 6 months ago is out of date!
I'm not sure if pcie3 will kill performance running more cards with models split over them?