← r/LocalLLaMA
▲
18
 
1👁
r/LocalLLaMA · u/CommonMinimum587 · 7h ago

3.66B formal logic model built on Granite 4.2 3B · Hugging Face

webAI released TwIL LM3 Pro on September 30. I haven't seen it posted here yet, so I went through the model card, and their comparison chart is attached.

It's a 3.66B model built on IBM's Granite 4.2 3B and tuned only for formal logic. That means things like checking whether a conclusion follows from its premises, rule induction, entailment and critiquing Lean proofs. They post trained it with LoRA SFT, checkpoint merging and RL against a programmatic verifier. The same recipe lifted VibeThinker-3B from 37.4 to 54.1 .

On their logic composite it scores 55.4, against 43.1 for the Granite base, 42.2 for the original TwIL-LM3, 41.2 for VibeThinker-3B and 53.4 for Qwen3-8B. The card itself calls the Qwen3-8B gap sampling noise, so that's a tie at less than half the size. Where it clearly leads is strict multiple choice logic, at 41% against 17% for the next model, and BBH logic at 95.4%.

They also publish where it loses. gpt oss 120b is still ahead on rule induction, entailment and Lean formalization. On general benchmarks it averages 79.0, against 81.0 for VibeThinker-3B and 84.9 for Qwen3-8B. It also thinks long on logic tasks, around 1,900 tokens per answer, and a quarter of answers hit the length cap

llama.cpp

If you want to try it, the Q4_K_M GGUF is 2.09 GiB and runs on CPU or 4 GB of VRAM, with Q5, Q6 and Q8 builds up to 3.63 GiB. It runs in Ollama, LM Studio and llama.cpp straight from the Hugging Face page, and the model card lists the exact commands. Keep the temperature at 0 to match their numbers, and give it at least 2048 tokens so the thinking doesn't get cut off. Their scores are on BF16 weights and webAI hasn't benchmarked Q4 yet. The license is non commercial

want to test it on policy rules with exceptions, contract conditions, and as a checker step in an agent pipeline before anything acts. If you've run it, how did Q4 hold up against their numbers, and did the long thinking get in the way?

Model card and full eval tables: https://huggingface.co/webAI-Official/TwIL-LM3-Pro

posted Sat, 10 Oct 2026 04:28:14 GMTseen 1 time
open on reddit ↗ 💬 5