An RL-tuned version of Qwen2.5 Coder on agentic tasks to edit codebases. The data was generated by using the model itself - successfully applied edits from one model is used to train the next iteration.