Open-Reasoner-Zero-32B18 Feb 2025
An open replication of DeepSeek-R1-Zero-Qwen-32B, i.e., a fine-tuned model of Qwen2.5 32B with RL only, missing the SFT cold start. They share the used code and data, as well as a tech report. This work will likely be covered more in a future post, but you should think of it as the most robust RL-on-base-model report since the DeepSeek R1 release.