All models

DeepSeek-R1-Distill-Qwen-32B

byDeepSeekDeepSeek· 20 Jan 2025
General purpose

Deepseek has released the full version of r1-lite, which we also covered in our blogpost. Aside from the official benchmark numbers, the model is remarkably strong and on par with o1, i.e., it currently is one of the best models available, both open and closed. The model is available under MIT, a license that DeepSeek has used for their experimental models like Janus in the past. Alongside the R1 and R1-Zero models, which are MoE, they also distilled the models with SFT onto Qwen and Llama of various sizes, with the appropriate licenses matching the original models. These models are incredibly strong for their size and continue the trend of distilling big models onto smaller ones, just like it has been done by the other, western labs for quite some time now. The distilled models are trained with SFT only - in the report, they mentioned that applying RL yields even further gains.

Specs
Params32B
LicenseMIT
Similarity · VAIL
VAIL
VAIL Fingerprint
003e:004f:006c:0098:00ee:0184:06dc:5f9c

Explore other models with behavioral similarity to DeepSeek-R1-Distill-Qwen-32B.

Capability · Artificial Analysis
AA Index
11.0
Months Behind Frontier
10.6 mo
claude-3-opus
04 Mar 2024
Adoption · Hugging Face
RAM score
Hugging Face Downloads
810.2K
last 30d
26.6M
all time
HF Likes
1.6K

Relative Adoption Metric contextualizes downloads against the model's size bucket.

Inference · OpenRouter
Tokens/Day
below top 50 this week
Peak Tokens/Day
620.3M
04 Feb 2025
Peak Rank
#22
04 Feb 2025

OpenRouter publishes daily token totals for its ~50 most-served models. Days missing from the chart mean DeepSeek-R1-Distill-Qwen-32B fell below that cutoff — not zero usage. OpenRouter logs usage separately per dated model version and per variant (like ":free"); the hub sums them into one daily total per model.

Related Models