I trained a 3.87B MoE (1.45B active) from scratch on only 86.5B tokens
First of all, thank you for reading.
I trained a small MoE model completely from scratch (no external base weights) and wanted to share the results + a couple of lessons.
Apex-2
\- Architecture: Decoder-only MoE, every layer is MoE (no dense layers)
\- Size: 3.87B total parameters, 1.45B active per token
\- 32 layers, d\_model 2048, GQA 16Q/4KV, 16 experts, top-4
\- Context: 4096
\- Tokenizer: Qwen3 (151k)
\- Hugging Face: https://huggingface.co/YOON1v/Apex-2
(loads with transformers / vLLM via Qwen3MoeForCausalLM mapping)
Training
\- Pretrain: 86.5B tokens (GH200 ×1 → ×2 with DiLoCo)
\- SFT: \~2.5B tokens (code-heavy + math + instruction)
\- DPO: tried it, scores dropped, so I dropped the checkpoint
Key numbers (SFT, greedy, chat template)
Benchmark
HumanEval 43.9
HumanEval+ 41.5
MBPP 56.3
MBPP+ 48.9
GSM8K (0-shot CoT) 32.4
MATH-500 21.0
IFEval (prompt strict) 44.7
MMLU (5-shot) 28.6
interesting comparison
With only \~0.087T pretrain tokens, the base model’s HumanEval+ matched Qwen2.5-1.5B (which used 18T).
Knowledge (MMLU) and math still lag far behind, as expected with the data gap.
What didn’t work
DPO (220k pairs, length-normalized) made answers much longer and hurt code / math / IFEval.
I stopped it and kept the SFT checkpoint. Full log is in the benchmark write-up.
Limitations (honest)
\- English-centric (almost no multilingual ability)
\- Weak knowledge → frequent hallucinations
\- LiveCodeBench medium/hard is near zero
\- 4k context only
Happy to answer questions about the MoE setup, DiLoCo, or why DPO backfired.