← r/LocalLLaMA
▲
20
 
1👁
r/LocalLLaMA · u/pseudotensor1234 · 4h ago

H2O-Lightning-4B: Apache-2.0 4B Decision model, official #1 open model on JevBench (above Jev)

Disclosure: I work at H2O.ai.
  
We released H2O-Lightning-4B, an open-weight (Apache-2.0) model for the "decisions API" style of inference that Jev made popular: you send a state plus typed questions (pick one / yes-no / score), and get calibrated probabilities back from a single forward pass. No generated tokens, so it's fast and cheap.
  
\*\*Results (JevBench, public leaderboard):\*\*
\- Composite score 72.5, vs Jev 1.13 at 71.5; currently the top open model
\- Leaderboard: https://benchmarkheaven.com/jev-models
  
\*\*Running it:\*\*
\- Base: Qwen3.5-4B, fine-tuned
\- Stock vLLM plus a small open shim (in the repo); \~30 ms per decision on an H100
\- Your data stays local, no per-call fees
  
\*\*Coming soon:\*\* 12B and 31B versions, which in our internal testing are considerably smarter than
 Jev, still open-weight and still one forward pass per decision.
  
 \*\*Demos\*\* (inbox triage of 1,000 insurance claims, a multi-browser web agent, DOOM on the decision clock): https://youtu.be/2Qp04Wu0A14
  
Weights, model card and serving instructions: https://huggingface.co/h2oai/h2o-lightning-4b
  
Happy to answer questions about the setup and latency.

posted Fri, 09 Oct 2026 20:26:12 GMTseen 1 time
open on reddit ↗ 💬 13