d1-3B and d1-omni from LiquidAI
https://preview.redd.it/owhrvbbiu2uh1.png?width=4096&format=png&auto=…
d1-omni-600M
d1-omni-600M is a 600M parameter decision model built on LFM2.5-Encoder-350M. You give it a state (text or JSON, with images or a voice clip) and a set of named questions. It returns typed answers with zero output tokens: every answer is read directly from the model's distribution over the options, with no generation and no parsing.
- Vision-language: text and images (tiled for large frames, several images per state) in a single forward pass.
- Audio-language: text and up to 30 s of speech in a single forward pass.
- Edge-sized: 587M parameters: a 381M shared trunk and decision head, a 94M vision encoder and a 112M audio encoder. Every modality runs the same trunk weights.
https://preview.redd.it/f2wkfqcku2uh1.png?width=1200&format=png&auto=…
d1-3B
d1-3B is a 3B parameter decision model built on LFM2.5-VL-3B. You give it a state (text, JSON, images, or a mix) and a set of questions. It returns calibrated, typed answers in one forward pass with zero output tokens.
- Best decision model under 10B on the Decision Index 0.2.1: 48.57, ahead of every 4B and 9B model and of Decider 35B-A3B (47.11).
- Multimodal: images and text in the same state. It scores 74.1 on 11 public image benchmarks (LFM2.5-VL-3B: 73.9).
- Fast: 8 ms a decision on an NVIDIA RTX 4090, 9 ms on an AMD MI325X, 30 ms on an Apple M5 Pro.
https://huggingface.co/LiquidAI/d1-3B-GGUF
https://huggingface.co/LiquidAI/d1-3B
https://huggingface.co/LiquidAI/d1-omni-600M-GGUF
https://huggingface.co/LiquidAI/d1-omni-600M
https://preview.redd.it/e11c4qlut2uh1.png?width=1932&format=png&auto=…