13 posts · 1 sub · RSS
← prev Friday, September 11, 2026 next →
posted hourdayweekmonthyearall
allr/LocalLLaMA
▲
123
+4
23👁
r/LocalLLaMA · u/running101 · 29d ago
nvidia rtx 5090 with 96gb of vram.

China-modified Nvidia RTX 5090 with massive 96GB of memory appears on Alibaba for less than $4,000 — 3x more VRAM at 65% the cost of the original

Anyone here running one of these? Or brave enough to purchase ?

Edit: I sent them an inquiry. They replied they can get me 5 x 5090 for $6k . Or some 4090 with 48gb .
I am going to keep messaging and questioning them. See where this goes.

Edit: so far they are denying having a 5090 96gb card. They offered a 48gb 4090 card. I am still discussing with them.

Edit: 9/14/2026: They quoted this, RTX 4090 48GB - 4286usd/pc
Still discussing with them

▲
155
-2
24👁
r/LocalLLaMA · u/Porespellar · 29d ago
Is a ZIMA Board 2 + RTX 2000 ADA the cheapest path to a decent Qwen-3.8 27b self-contained endpoint? post image

I just watched a YouTube from Luke’s Dev Lab where he literally just plugged a RTX 2000 ADA Into the side of the Zima Board 2’s PCIE socket and it just friggin worked and had great token speed despite running on shitty Ollama. Ran off the Zima’s power supply and everything.

https://youtu.be/Lb3sRFTA-hk?si=8S8vv4GD1zVPeTrc

The Zima Board 2 is only like $411. It has like 16GB RAM and 64 GB eemc storage, Sata ports, Ethernet, yada, yada.

https://shop.zimaspace.com/products/zimaboard2-single-board-server

an Nvidia RTX 2000 ADA is like $700 and has 16GB of VRAM. $1100 for both seems like a great entry point for having a fully functional Qwen 3.8 27b endpoint running at a decent tk/s.

Is this the cheapest and best-performing self-contained entry point for local AI or would a baseline (pre order) Mac Mini M5 with 24GB be a better way forward. Seems like the RTX would still edge out the M5 Mac for prompt processing speed but you do get a much better actual computer in the Mac.

Are there any cheaper fully self-contained alternatives that offer fast token speed on a decent size model like Qwen 3.8 27b?

I’m focusing the discussion on new systems you can buy or preorder now and not used systems. I’m sure there are great deals on used Macs out there, but I want good prefill speeds.

▲
62
 
28👁
r/LocalLLaMA · u/MountainTop321 · 29d ago
CodeFinetuner: Fine-tune a local code autocomplete model on your own codebase post image

Hi everyone,

I was interested in learning LoRA fine-tuning, and ended up building CodeFinetuner over the past few months, a full pipeline that fine-tunes a small code autocomplete model (e.g. Qwen2.5-Coder-3B) specific to a codebase. You can then use the resulting GGUF model via llama.vim/llama.vscode and run it fully locally. Supports fine-tuning on Mac (MPS) and NVIDIA GPUs (CUDA), with optional Unsloth support for faster training and lower VRAM usage.

Pipeline: raw code -> tree-sitter parsing into Structure-Aware FIM examples -> LoRA fine-tuning -> evaluation (CodeBLEU, edit similarity, exact match, perplexity, ...) -> GGUF conversion for local inference.

To try it:

uv tool install codefinetuner

Create a data folder and place your repo (or code files) inside. For auto-split just drop the files in directly, for manual split create data/train/, data/eval/, data/test/ subfolders and set split_mode: "manual". Get the default config with:

curl -L -O https://raw.githubusercontent.com/cuolm/codefinetuner/master/config/codefinet…

Adjust it to your needs and hardware availability, then run:

codefinetuner --config="codefinetuner_config.yaml"

The example runs in the repo show clear improvements over the base model on these evaluation metrics, but using the model for autocomplete on code you're actively writing is a different thing from scoring well on a test set, and the autocomplete tools themselves (llama.vim/llama.vscode) sample differently from the greedy decoding used in the evaluation. So the real usefulness still has to be verified in the editor itself.

Might also be useful just as a reference, since it's a complete working LoRA fine-tuning pipeline end to end.

Hope someone finds this project interesting or helpful.

https://github.com/cuolm/codefinetuner

▲
94
+2
28👁
r/LocalLLaMA · u/jqwl · 29d ago
Any 12gb VRAM users out there?

Hi!

I've been following this community for quite a while and have difficulty figuring out what to put on my 3080 12gb - I know Qwen 3.6 35B 3A was the go-to choice when it first came out, but I'm curious if there are any other models / specifically optimized models that meaningfully benefit from the extra 4gb of VRAM over 8gb while still being usable under 16gb.

My workflow is agent heavy, but more for a personal secretary and manager, and less coding heavy.

Thanks!

💬 90 (+1) open on reddit ↗
▲
760
+5
33👁
r/LocalLLaMA · u/kvyb · 29d ago
Qwen3.8-27B-Humanlike-Chat: A model I tuned to imitate realistic human-to-human conversation post image

I made this because I was getting genuinely annoyed at trying to have a normal conversation with LLMs. Even with prompting and various tricks, most models I've tried still have this "AI assistant" vibe to them that is so familiar: too helpful, polished, verbose, using words we never use in conversation, etc.

I wanted a model that could just talk to me like a person, so I did the slightly unreasonable thing and put together a dataset and trained one.

The dataset used for training is 125,217 obfuscated human-to-human messages across 1396 chat conversations.

The goal wasn't to make Qwen smarter or improve benchmark scores. I was trying to change its conversational habits, to make it stop turning every reply into an explanation, agreeing with everything, and writing stuff just to keep the conversation "going".

I trained a rank-256 LoRA on top of huihui-ai/Huihui-Qwen3.8-27B-abliterated. The released version is checkpoint 863. In my testing it feels noticeably less like an assistant, particularly in casual conversations, even without a system prompt. Replies are generally shorter, less polished, and, well, more human.

There may be a tradeoff. An earlier iteration scored five percentage points lower than its Huihui parent on IFEval, an instruction-following benchmark. I haven't rerun that benchmark on this version of the checkpoint, and I haven't tested coding performance, so I don't want to pretend that number applies here.

I've added a side-by-side comparison using the same system prompt, user messages, and generation settings for both models. Each model continued its own conversation branch, with reasoning effort set to 'xhigh'.

Merged GGUFs and the standalone F32 LoRA adapter are in the model repo:

https://huggingface.co/LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-GGUF

Space where you can have a demo chat with different system prompts and reasoning modes:

https://huggingface.co/spaces/LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat

UPD: I certainly didn't expect this post to blow up like this! There's been a lot of great discussion in this thread and a lot of insight for me on where to take the model next.

A few have asked for our Discord, and we'd be happy to see you there: https://discord.gg/aCCrWftMjS

▲
99
-1
16👁
r/LocalLLaMA · u/arturdent · 29d ago
Orukeet, new ASR model based on Parakeet

I haven't seen this mentioned yet, so I thought it deserves a post. I was trying out OpenWhispr when this model came up as the recommendation. So I don't have personal experience yet, but it's supposed to be a better version of Parakeet, especially on Macs.

Their official tidbit:
"Orukeet is a 25-language speech recognizer built from NVIDIA Parakeet TDT 0.6B v3. It replaces half of the encoder's temporal depthwise filters with 12,288 fitted, frozen Gabor kernels and trains the remaining parameters on multilingual and multi-accent data.

Orukeet outperforms Parakeet on 61 of 74 tested splits, including LibriSpeech test-clean (1.46% vs. 1.53% WER), test-other (2.86% vs. 3.14%), and FLEURS English (3.82% vs. 4.28%). Across all 25 FLEURS languages, pooled WER is 9.85% vs. 11.01%, a 10.6% relative reduction. Final adaptation and checkpoint selection use LibriSpeech test-other."

https://huggingface.co/oruk/orukeet

▲
165
-4
26👁
r/LocalLLaMA · u/Ok_Warning2146 · 29d ago
Terminal Bench v4 scores post image

Some people says terminal bench reflects model intelligence better than the intelligent index. From the look of it, the ranking does seem to reflect how people feel about the open and closed models.

For the open models, GLM-5.3 is in a league of its own. GLM-5.3-Flash is leading the current gen of top flash models. Kimi-K3 did pretty bad in this benchmark for its size. Qwen3.8-27B is the only small model that can do something on this bench.

|Model|Score|
|:-|:-|
|GLM-5.3|41.9%|
|GLM-5.3-Flash|32.8%|
|DSV4.1-Flash|26.8%|
|Qwen3.8-Flash-Next|25.3%|
|DSV4-Pro|14.1%|
|Kimi-K3|12.6%|
|DSV4-Flash|12.1%|
|Qwen3.8-27B|5.6%|
|Muse Glimmer|0.5%|
|gemma4-31b|0.0%|

▲
3068
+69
64👁
▲
76
+3
20👁
▲
66
-3
31👁
r/LocalLLaMA · u/pmttyji · 29d ago
CUDA/HIP: Flash Attention tuning (gfx1201) by pwilkin · Pull Request #28102 · ggml-org/llama.cpp

Nice pp improvements for RDNA4(R9700) & 3.5(RX 9060 XT, 8060S). More good numbers on large context.

PR has detailed benchmarks.

u/ilintar 👍

▲
64
+2
35👁
r/LocalLLaMA · u/riceinmybelly · 29d ago
What can you run on 8GB VRAM?

Can you still do something with a 2050 or something like it?
I mean for office work, loading embedding, reranking and chat models not at the same time but is anyone still using smaller models and have any good ones come out?

I feel like small models are abandoned, I don’t care much for world knowledge, I want tool use and preferably multilingual. Vision would be nice but beggars can’t be choosers.

▲
542
+2
28👁
r/LocalLLaMA · u/T_rex2700 · 30d ago
Someone apparently managed to kind of replicate what V4.1 flash does on KV for fast prefill on Qwen post image

I wonder someone will figure out a way to do this with 27B?

Throw Qwen3 on this page for demo
https://kishida.github.io/webdemos/llkvapprox/

Edit: sources (thank you u/pmttyji for finding them!

▲
407
+7
35👁
r/LocalLLaMA · u/Acceptable-Cycle4645 · 30d ago
New Music Model YuE2-3B Released!

Surprised no one has posted it in this sub.

Pretty solid model, IMHO.

Demo: https://map-yue2.github.io/