s1.1-32B08 Feb 2025
A fine-tuned version of Qwen2.5 32B trained on just 1,000 examples. This version uses R1-generated traces, while the first version uses traces generated by Gemini Flash Thinking. However, the authors found that the R1 traces are longer and more diverse than the Gemini traces.