Aplomb 1 is #1 on Typed Decisions
Aplomb 1 is #1 on Typed Decisions, an official Hugging Face benchmark for probabilistic decisions. At 5.3B parameters, its probabilities are the closest to the true answers on the leaderboard, by KL from gold and by Brier score, ahead of a 27B model about five times its size.
Typed Decisions gives a model one piece of unstructured text and five typed questions about it at once. Every answer is a probability distribution, scored against the gold distribution across 400 cases and 2,000 decisions. KL from gold and the Brier score measure how far a model's probabilities sit from the true answers, so lower is better.
Results:
- KL from gold: 0.123, #1 on the leaderboard and 40% lower than the 27B model in second place
- Brier score: 0.065, #1 on the leaderboard and a third lower than the 27B model in second place
Aplomb 1 takes text, JSON, images, video and audio in one request and reads up to 1M tokens. On our API it answers a short question in about 15 ms of model time and reads a 1M-token document in about 3 seconds, at $0.02 per 1M input tokens with output free. The weights are open on Hugging Face.
Leaderboard: https://huggingface.co/datasets/LocalLLaMA/typed-decisions?leaderboard\_task\_id=kl\_from\_gold