DeepSeek

DeepSeek

22 models in Artifacts Logshuggingface.co/deepseek-ai
DeepSeek-V4-Flash-073131 Jul 2026
284B (13B active)MIT156.2K · 30d

Just one day after OpenAI has dropped the prices of their smallest model by 80%, the whale dropped an update to their V4 Flash model, beating Luna at the pareto frontier. The bigger model is not updated yet, so it remains to be seen where it will land in terms of performance. For the initial V4 releases, the Flash version was the star of the show in terms of performance per parameter, while Pro was rather underwhelming.

DeepSeek-V4-Flash24 Apr 2026
284B (13B active)MIT2.8M · 30d

DeepSeek has finally released its successor to the V3 series, which it kept updating for months. It comes in two sizes: Pro, which is a 1.6T-A49B MoE, and Flash, a 284B-13B model. Based on others' experience, the latter model seems to be the real star of the show, as its performance is relatively strong, while Pro seems to underdeliver relative to its size. The tech report goes into great detail, including the architectural changes used to achieve better and cheaper long-context performance.

DeepSeek-OCR-227 Jan 2026
3.4BApache-2.02.2M · 30d

An updated version of DeepSeek-OCR.

DeepSeek-V3.201 Dec 2025
671BMIT1.2M · 30d

Another update to the V3 series, which integrates DSA. They also trained and released a "high compute" version, V3.2 Speciale, which claims to beat the 2025 IMO and IOI with gold-medal performance.

DeepSeek-OCR17 Oct 2025
3.3BMIT2.4M · 30d

A rather surprising release by DeepSeek showcasing the superior compression ratio of images to text as the main argument in their paper.

DeepSeek-V3.2-Exp29 Sep 2025
671BMIT258.6K · 30d

DeepSeek added a new sparse attention variant to their arsenal, showing impressive long context efficiency. This method works on existing models, surpassing the long context scores compared to previous versions.

DeepSeek-V3.121 Aug 2025
671BMIT283.4K · 30d

To the surprise of everyone, DeepSeek launched another iteration in their V3 series. This model combines the reasoning of R1 with the non-reasoning component of V3, which is different from Qwen which is splitting its models again.

DeepSeek-R1-0528-Qwen3-8B29 May 2025
8BMIT1.8M · 30d

DeepSeek updated their R1 reasoning model. The model card has not a lot of details (more compute + algorithmic improvements), but the benchmarks show that the whale is yet again at the frontier, rivaling closed models. Even more exciting as we enter the agentic era is the support for tools. R1 is close to the frontier in popular function-calling benchmarks as well: In my quick tests it seems like the new R1 at least has the same distinctive character, but could be even more clever. Here's an example: Another example is it seeming to do better on the search APIs (at least what is built into OpenRouter), boosted by the new tool-calling capabilities. This new version of R1 is much less likely to start every reasoning chain with "Okay," which should be good for diversity and token efficiency. And, to top it all off: They also used the new R1 thinking traces to fine-tune DeepSeek-R1-0528-Qwen3-8B, boosting scores significantly in math benchmarks. For more, see DeepSeek's announcement or Artificial Analysis's scores.

DeepSeek-R1-052828 May 2025
671BMIT211.2K · 30d

DeepSeek updated their R1 reasoning model. The model card has not a lot of details (more compute + algorithmic improvements), but the benchmarks show that the whale is yet again at the frontier, rivaling closed models. Even more exciting as we enter the agentic era is the support for tools. R1 is close to the frontier in popular function-calling benchmarks as well: In my quick tests it seems like the new R1 at least has the same distinctive character, but could be even more clever. Here's an example: Another example is it seeming to do better on the search APIs (at least what is built into OpenRouter), boosted by the new tool-calling capabilities. This new version of R1 is much less likely to start every reasoning chain with "Okay," which should be good for diversity and token efficiency. And, to top it all off: They also used the new R1 thinking traces to fine-tune DeepSeek-R1-0528-Qwen3-8B, boosting scores significantly in math benchmarks. For more, see DeepSeek's announcement or Artificial Analysis's scores.

DeepSeek-Prover-V2-671B30 Apr 2025
671BDeepSeek Model License716 · 30d

An updated version of the Prover series by DeepSeek, building upon V3. They also used a smaller 7B Prover model (which, of course, is also open-sourced) to search for subgoals in proofs. The resulting model shatters proof-related benchmarks, but this is expected from DeepSeek.

DeepSeek-V3-032425 Mar 2025
671BMIT912.8K · 30d

An updated version of DeepSeek V3, focusing on tool-use, reasoning and front-end development. In WebDev Arena, the updated version improved by a whopping 240 points. Even more exciting is the license change: The initial version was licensed under the DeepSeek license, this version is licensed under MIT, just like R1.

Janus-Pro-7B26 Jan 2025
7BDeepSeek Model License12.8K · 30d

Deepseek released a new version of their experimental omni-modal Janus models. Similar to the previous release, it is not ready for prime-time yet.

DeepSeek-R120 Jan 2025
671BMIT9.6M · 30d

Deepseek has released the full version of r1-lite, which we also covered in our blogpost. Aside from the official benchmark numbers, the model is remarkably strong and on par with o1, i.e., it currently is one of the best models available, both open and closed. The model is available under MIT, a license that DeepSeek has used for their experimental models like Janus in the past. Alongside the R1 and R1-Zero models, which are MoE, they also distilled the models with SFT onto Qwen and Llama of various sizes, with the appropriate licenses matching the original models. These models are incredibly strong for their size and continue the trend of distilling big models onto smaller ones, just like it has been done by the other, western labs for quite some time now. The distilled models are trained with SFT only - in the report, they mentioned that applying RL yields even further gains.

DeepSeek-R1-Distill-Qwen-32B20 Jan 2025
32BMIT810.2K · 30d

Deepseek has released the full version of r1-lite, which we also covered in our blogpost. Aside from the official benchmark numbers, the model is remarkably strong and on par with o1, i.e., it currently is one of the best models available, both open and closed. The model is available under MIT, a license that DeepSeek has used for their experimental models like Janus in the past. Alongside the R1 and R1-Zero models, which are MoE, they also distilled the models with SFT onto Qwen and Llama of various sizes, with the appropriate licenses matching the original models. These models are incredibly strong for their size and continue the trend of distilling big models onto smaller ones, just like it has been done by the other, western labs for quite some time now. The distilled models are trained with SFT only - in the report, they mentioned that applying RL yields even further gains.

DeepSeek-V326 Dec 2024
671BCustom Model Agreement1.2M · 30d

A new release by DeepSeek, which we covered in our post. It undeniably is among the strongest models out there, which cannot be ignored: Neoclouds, such as TogetherAI or DeepInfra are starting to host the model despite its big size requiring two H100 nodes.

DeepSeek-VL213 Dec 2024
1.0B, 2.8B, 4.5BDeepSeek Model License2.8K · 30d

DeepSeek has released a new version of their vision models. The models come in various small sizes and are all MoE models, which is unusual in this size range. The language models are trained solely for this model, which could be from experimentation when training DeepSeek V3.

Janus-1.3B18 Oct 2024
2.1K · 30d

Deepseek released a new codebase, paper, and model, Janus-1.3B, which is a new image+text input and output model. This seems to cover everything - new architecture, data, and fine-tuning practices. Note, that these outputs are not interleaved, so it is either text or image outputs in blocks.

DeepSeek-V2.505 Sep 2024
5.7K · 30d

The most interesting part of this chat model is that they merged two post-training paths, chat, and math, to result in one model better than the sum of its parts!

DeepSeek-V2-Chat-062818 Jul 2024
4.8K · 30d

The top open-weight model on ChatBotArena before Llama 3.1 was released. They even beat Gemma 2 27B!

DeepSeek-Coder-V2-Instruct14 Jun 2024
7.1K · 30d

A super popular new coding model. Evals on coding specific models like this are tending to match or pass the API-based general models. I haven't given them a shot yet.

DeepSeek-V2-Lite15 May 2024
494.2K · 30d

Another great chat model from Chinese open model contributors. Consistently, the 01-ai, DeepSeek, and Qwen teams are shipping great models This DeepSeek model has "16B total params, 2.4B active params" and is trained on 5.7 trillion tokens. This is a great size for many people to play with.

DeepSeek-V222 Apr 2024
10.2K · 30d

Another strong MoE base model from the DeepSeek team. Some people are questioning the very high MMLU scores, which is a similar story to the next model (Yi 1.5). Regardless, open models are making a lot of progress on MoE models. Scaling MoE models from this 20B active range to 100+ is supposedly an almighty engineering challenge.