FreedomIntelligence/HuatuoGPT-3-27B · Hugging Face
from FreedomIntelligence: HuatuoGPT-3-27B is a medical LLM built on Qwen3.8-27B with One-stage Policy Optimization (OnePO). OnePO adapts language models to medicine in a single reinforcement-learning stage, without preceding domain-specific supervised fine-tuning. Teacher responses provide temporary guidance and are retired as the model improves. We release the training code, medical RL dataset, and 8B rubric grader. (last week they released https://huggingface.co/FreedomIntelligence/HuatuoGPT-3-9B)
scorecomments7 sightings
first seen 2026-10-03 06:32 UTClast seen 2026-10-03 13:39 UTCscore then 39score now 40gained +1sightings 7
open on reddit ↗
💬 3