LiquidAI/LFM2.5-Encoder 250M/350M
#
LFM2.5-Encoder-350M is a multilingual bidirectional encoder built on the LFM2 architecture — a larger encoder for maximum downstream quality. It is a masked language model with full bidirectional attention, designed to be fine-tuned into task-specific models (classification, token classification, retrieval, reranking, and semantic similarity) across 15 languages, and to run efficiently on-device.
- Highly capable for its size. On par with the best similarly sized encoders and well ahead of our own retrieval siblings.
- General-purpose. 8k context, strong across NLI, paraphrase, sentiment, and multilingual tasks.
- Fast and on-device. Matches or beats ModernBERT throughput, with a long-context edge on CPU; runs in the browser on WebGPU.
https://huggingface.co/LiquidAI/LFM2.5-Encoder-350M-GGUF
https://huggingface.co/LiquidAI/LFM2.5-Encoder-230M-GGUF
https://github.com/ggml-org/llama.cpp/pull/29862
example (from the hf):
❯ uv run fill-mask.py LFM2.5-Encoder-350M-F16.gguf "The capital of France is [MASK]."
top-5 at [MASK]:
# 1 11.42 ' Paris'
# 2 10.43 'Paris'
# 3 9.65 ' Nice'
# 4 8.94 ' Strasbourg'
# 5 8.62 ' Lyon'