
Agentica
2 models in Artifacts Logshuggingface.co/agentica-org
DeepCoder-14B-Preview07 Apr 2025
14BMIT788 · 30d
A RL-trained version of the R1-version of Qwen2.5 14B. They use the insights from DAPO for an improved version of GRPO, which we covered here.
DeepScaleR-1.5B-Preview29 Jan 2025
1.5BMIT9.6K · 30d
A RL-tuned version of R1-Qwen2.5 1.5B, which improves the math performance compared to the original, distilled model significantly. One of their main findings is that the incorrect responses were three times longer than the correct ones.