4090 48G +128G+strata test
Measured on a 4090 48GB + 128 GB RAM: keeping Strata's 28.8 GB n-gram table in RAM buys \~1%, while conversation parking bought me 35x
Alternative: I benchmarked 4 ways of placing Strata's n-gram table. The default is already right - here's what actually moved
Body
Everything below is measured on one machine. Where I don't have a number, I don't make a claim.
TL;DR (all measured)
- Whole n-gram table in RAM: +0.65% prefill / +1.2% decode, cost +28.4 GiB RAM. Not worth it.
--ple-io mmap: 72.8 s vs 19.5 s on the first long prompt (3.7x slower), and it silently disables--ple-row-cache.- My earlier "+4.7%" for the RAM-resident table was a wrong baseline, not a real effect. Session-to-session spread on an unchanged config was 11.5%.
- Conversation parking: 18.8 s → 537 ms to return to a 91,836-token conversation, for 3.1 GB of RAM.
--prefill auto:32768: +9.7% prefill, +13.6% decode at a 78.7K prompt.--calibratethen found +3.2% by changing one value I'd never have guessed (23 → 12 CPU workers).
Setup
|||
|:-|:-|
|CPU / GPU / RAM|i9-13900K / RTX 4090 48 GB (driver 617.14) / 128 GB DDR5-6000|
|OS / engine|Windows 11, Strata 0.1.38 release build (sm\_89), IQ3\_S|
|Config|262K context, --kv int8 --kv-resident 32768, --expert-cache auto --prefill auto --spec 4, vision on|
|Model files|shard1 54.8 GB + shard2 28.8 GB, both sha256 == published values|
Engine log at startup (relevant later): experts loaded: 46.84 GiB at 5.12 GiB/s (19 s) and expert cache auto: 40.69 GiB free, 700 MiB reserved (+218 MiB draft) -> 20880 slots, 39.79 GiB of VRAM — i.e. 20,880 of 24,576 experts (85%) in VRAM, decode hit rate 97.3%-99.3%.
Baseline throughput on this box: decode 128-151 tok/s, prefill 4,249-5,154 tok/s at 78K-92K context (measured with my own harness and with lm-eval-harness).
1. Where the 28.8 GB n-gram table lives — 4 arms
Same 91,836-token real-text prompt, fresh engine start per arm, same benchmark script, one run per arm:
|Arm|--ple-io|--ple-row-cache|RAM used|Cold prefill (disk read / tok/s)|Warm prefill (disk read / tok/s)|Decode|
|:-|:-|:-|:-|:-|:-|:-|
|A|direct|1,048,576 rows (\~90 MB)|68.2 GiB|3,114 MB / 4,872|242 MB / 5,201|128.7|
|B|mmap|same|68.0 GiB|72.8 s / 1,274|0.4 MB / 5,208|99.8|
|C|direct|whole table (320,001,536 rows)|96.6 GiB|3,111 MB / 4,872|1.4 MB / 5,235|130.3|
|D|mmap|whole table|68.4 GiB|71.6 s / 1,295|0.4 MB / 5,193|127.9|
Conclusions:
- A vs C is the only clean comparison (same mode, only cache size): +0.65% warm prefill, +1.2% decode. C reads 173x less from disk and runs essentially the same speed → the n-gram table on the SSD is not a bottleneck on this machine.
- The default \~90 MB row cache already absorbs 92% of the reusable traffic (3,114 MB → 242 MB on the second pass over the same text).
mmapis worse: the first long prompt is 3.7x slower (cold page cache), and arm D only used 68.4 GiB RAM — the 28.8 GB was never allocated, so the row cache is a no-op undermmap. (The "0.4 MB read" in B/D is an artifact: mmap faults don't appear in the process read counters.)
My own mistake, worth repeating: my first pass reported +4.7% for C. It came from a different session than the baseline. Later, with an unchanged config, I measured 149.1 → 166.3 tok/s (11.5% spread) between sessions on identical settings. If you're A/B-ing anything here, run both arms back to back in the same session — otherwise you publish noise.
2. Conversation parking — the biggest effect I measured (35x)
Added "--conversation-cache-mib", "8192" \+ "--conversation-cache-slots", "4", then alternated two unrelated long conversations:
|Step|Wall clock|Disk read|Engine log|
|:-|:-|:-|:-|
|P1 first time|19.5 s|2,969 MB|91836 tokens = 0 reused + 91836 read|
|P2 (other conversation)|17.2 s|614 MB|P1 parked: 91,870 tokens / 493 ms / 1.88 GB|
|Back to P1|1.2 s|1.5 MB|91829 reused + 7 read in 537 ms|
RAM cost 3.1 GB of an 8 GiB budget, no evictions. 28.4 GiB bought 1%; 3.1 GiB bought 35x.
3. --prefill auto:32768
Before, the log said prompt chunk auto: 8192 tokens; after, 32768. Measured at a 78.7K prompt: prefill 4,667 → 5,106-5,136 tok/s, decode inside that context 112.6 → 127.2-129.1 tok/s. Short prompts didn't move, so if you test this with a short prompt you'll conclude it does nothing.
4. --calibrate — don't hand-tune
代码块
PCIe share 0.00 -> 150.8 | 0.20 -> 153.0 | 0.35 -> 154.1 | 0.55 -> 155.0 | 0.75 -> 146.9
draft floor 0.30 -> 140.2 | 0.50 -> 143.4 | 0.70 -> 140.1
CPU workers 23 -> 148.0 | 15 -> 149.9 | 12 -> 152.7 <- picked
It changed one value, --pool-workers 23 → 12 (+3.2%), on a 13900K. Verify the file afterwards — it prints the summary even if the JSON edit failed (it failed once on me). Re-run it after any config change: I did, and it picked 12 again.
5. expert_profile_save — works, no measurable gain
Verified the whole chain: engine counts routing → POST /unload writes a 192 KB profile → on the next server start the config's --expert-profile is replaced by the learned file (confirmed from the actual process command line). Measured effect: long prefill 5,059 vs 5,120 tok/s, long decode 118.4 vs 127.9 — i.e. inside noise, with hit rates already at 98-99%. Cost is 192 KB and no VRAM, so I left it on, but it is not a speed win.
6. Smaller measured things that cost me time
- 256K context doesn't tax short chats: 17-token prompt decodes at 128-133 tok/s; the 78.7K prompt decodes at 112.6 and reads 110 MiB of KV from RAM (vs 0.6 MiB). Cost tracks what you use, not the configured ceiling.
- Vision works and is cheap-ish: synthetic test image read in 0.97 s and described correctly; cost is 641 expert slots (-3.1% of the cache), text throughput slightly lower (128-133 vs 136-151 short / 112.6 vs 117.1 long).
- Sharing the GPU with ComfyUI:
POST /unloadfrees all 47.9 GB, andPOST /load\+ first token back took 18-19 s. That's what I use now. - Do not minimize the engine's console window on Windows 11: 45.8 → 36.6 tok/s (-20%), and the engine's own log shows it (
running on E-cores (EcoQoS background mode), -20%). Covering it with another window is fine. I checked 0.1.38: the fix isn't in it. - PowerShell 5.1 mangles non-ASCII request bodies (
Invoke-RestMethodsends ISO-8859-1, so the model literally sees?????). Usecurl --data-binary @fileor raw bytes from Python — same bytes, correct answer. - Manually placed model files need a
<filename>.donemarker, or setup re-downloads (I wasted a 0.9 GB download). - A
--setuppass rewrites the config and drops hand-added args. After one, my measured best settings were gone (chunk back toauto, workers back to 23) and throughput fell from 151.6 to 133.0-133.3 tok/s. Diff the config after every setup run. - Slow Hugging Face from CN: 87 KB/s through my proxy vs ModelScope at 12.3 MB/s per connection, 92 MB/s with 7 parallel streams (83.6 GB in \~12 min).
HF_ENDPOINT=https://modelscope.cn/modelsworked for model, MTP and mmproj. - Three env flags some contributor branches document (
STRATA_PREFILL_OWN_AUTO,STRATA_IQ256_GATHER,STRATA_KV_PREFETCH) are not in the 0.1.38 prebuilt binary — checked by reading the binary's strings. Setting them does nothing without your own build.
What I did not test
Quality of any kind (no perplexity, no KL, no task suite), other GPUs or AMD, a second GPU, --kv q4_0, contexts past 262K, and slower storage than my NVMe — so I can't say when the n-gram table would matter, only that it doesn't here. Some decode samples in section 1 are only 45 tokens long, which is why I put no weight on the 1.2% difference in arm C.
Repro: I switched arms by editing those two keys in strata-<model>.json, restarting the server, polling /health until loaded: true, then sending the same 4 requests in the same order. Scripts in the comments if useful.