A "true" distillation of DeepSeek V3 onto Qwen2.5-32B. Most models use supervised fine-tuning on outputs generated by the teacher, but this model is trained directly on V3's logits. For this to work, the model had to be re-trained on the V3 tokenizer, then distilled to finally use the original tokenizer and post-train the model.
Specs
Params32B
LicenseApache-2.0
Tags
Adoption · Hugging Face
RAM score
—
Hugging Face Downloads
48
last 30d
9.2K
all time
HF Likes
58
Relative Adoption Metric contextualizes downloads against the model's size bucket.
Related Models




