A series of reasoning models trained with SFT, followed by DPO and a GRPO variant as well as SEMPER during the RL phase. It performs surprisingly well given the low number of examples (10K) during the final RL phase.
Specs
Params32B
LicenseEXAONE AI Model License Agreement 1.1 - NC
Adoption · Hugging Face
RAM score
Relative Adoption Metric not scored because required parameter, download, or API metadata is not cataloged. This is not a zero score.
Hugging Face Downloads
470
last 30d
1.4M
all time
HF Likes
301
Relative Adoption Metric not scored because required parameter, download, or API metadata is not cataloged. This is not a zero score.
Related Models




