PSA: if you're on an Intel hybrid CPU, run Strata's calibrate - it nearly tripled my decode speed (IQ3_S at 256K, 16 GB card)
I polished it with GLM and it kinda sounds like AI. First time in the community, I used AI to polish it, but the AI copy is too wordy, so I sincerely apologize to you all... (sorry. This is the third version. In the third version, I added P-core thread pinning.) https://preview.redd.it/1035zf786hth1.png?width=852&format=png&auto=w… setup: 5070 ti 16gb, 96gb ram, i7-14700kf, windows. qwen3.8-flash-next iq3\_s on strata, 262k context.first test: \~17 tok/s at 256k. log screenshot attached, before lines are stock settings.then i changed 3 things: pool workers 13 instead of 19 (e-cores were stalling every verify window on my 14700kf), spec 6 + spec-min-p 0.7, pcie-frac 0. all measured by the built in calibrator, i didn't hand tune anything.now 256k sits around 55 tok/s warm (prefix cached, thats how agent sessions actually run). cold is 43.if you're on a 12th-14th gen intel cpu just run the calibrator, the defaults were measured on a 6-core ryzen with no e-cores.full numbers: https://github.com/JiuYue0820/Strata/blob/docs-256k-tuning/docs/TUNING-256K-16GB.md original text: 拿GLM润色了一下有点像AI,第一次来社区我用了AI润色但是AI文案太几把咯嗦了所以我像你们郑重道歉...对不起 然后就是这个是第三版,我由评论测了一下绑P核 我电脑配置是5070 Ti 16GB 显存,96GB 内存,i7-14700KF,Windows 系统 然后用的模型是 qwen3.8-flash-next iq3\_s,跑在 Strata 上,上下文 262K 在啥也没测试的时候256K 上下文下约 17 tok/s。日志截图在附件里,改动前的数据都是默认设置 我改了三个地方pool workers 从 19 改成 13(我的 14700KF 上,E 核在每个验证窗口都会造成卡顿)spec 设为 6 + spec-min-p 设为 0.7pcie-frac 设为 0全部用内置的校准器测得,我没有手动调任何参数。 现在 256k 在 warm 时大约 55 tok/s(prefix 已缓存,那就是 agent 会话实际运行的方式)cold 是 43 如果你在 12 代-14 代 Intel CPU 上,就运行 calibrator,默认值是在一个没有 E 核的 6 核 Ryzen 上测得的 完整数据:https://github.com/JiuYue0820/Strata/blob/docs-256k-tuning/docs/TUNING-256K-16GB.md