Gevva0 - a Jev like decision engine on Gemma 26B via direct logit scoring
On the official JevBench evaluation battery, Gevva0 scored 74.63 (#1 global rank), averaging 214ms p50 across large legal contract sets with 82.9% accuracy on the forensic hard tier.
How it works under the hood:
- Direct Logit Scoring: Ingests context and reads decision logits directly from llm.scores\[-1\] in a single prefill pass. Fast-path resolution runs in 22ms on short contexts.
- Cyclic Debiasing: Permutes class tokens across 4 cyclic positions to neutralize label position bias.
- Platt Temperature Calibration: Fits confidence via sigmoid scaling to push Expected Calibration Error (ECE) below 0.03.
- Asymmetric Audit Pass: Locks the categorical verdict permanently first, then performs an isolated extraction pass to retrieve verbatim source quotes without contaminating the decision logit.
The repo includes the evaluation harness, raw benchmark datasets, and a local web dashboard: https://github.com/solvingSteve/Gevva0
Setup instructions and benchmarks are in the README.
Working on Demos and Use Cases now so if you have any ideas I'll try to build them next!
scorecomments8 sightings
first seen 2026-10-03 06:30 UTClast seen 2026-10-07 08:02 UTCscore then 0score now 0gained 0sightings 8
open on reddit ↗
💬 12