BAAI/AREX-2 - 27B - Agent model based on Qwen3.8 27B
"AREX-2 is a 27B-parameter long-horizon agent model from the Beijing Academy of Artificial Intelligence (BAAI). It learns to improve a solution over multiple test-time rounds: propose, measure, reflect, and revise.
AREX-2 is trained on machine-learning and algorithmic-programming tasks with verifiable feedback, together with the existing AREX deep-research data. The learned self-improvement behavior transfers to deep research without adding new search trajectories.
- Architecture: Dense Qwen3.8-compatible multimodal model
- Parameters: 27B
- Context length: 262,144 tokens
[](https://huggingface.co/BAAI/AREX-2#key-features)Key features
- Long-horizon self-improvement: turns extra test-time rounds into useful solution refinement.
- Feedback-driven reflection: reads scores, logs, errors, and timings to decide what to change next.
- Cross-domain performance: training on coding and machine-learning tasks also improves the model's deep-research performance.
- Long-horizon reasoning: sustains productive iteration as the task budget grows."
scorecomments34 sightings
first seen 2026-10-03 04:48 UTClast seen 2026-10-09 01:16 UTCscore then 75score now 82gained +7sightings 34
open on reddit ↗
💬 29 (+2)