Inclusion AI

Inclusion AI

18 models in Artifacts Logshuggingface.co/inclusionAI
LLaDA2.2-flash16 Jul 2026
100BApache-2.0811 · 30d

An update to the diffusion model series from Ant Group/InclusionAI. Diffusion models are exciting for local use, as personal usage usually involves small batch sizes, a setting where diffusion models really shine.

Ling-2.6-1T29 Apr 2026
1TMIT1.5K · 30d

An update to the 1T MoE by Ant Ling.

Ling-2.5-1T14 Feb 2026
1TMIT643 · 30d

Ant Ling has not forgotten how to scale and has updated its 1T model, boosting scores across the board.

Ming-flash-omni 2.011 Feb 2026
100B (6B active)MIT2.7K · 30d

A 100B-A6B multi-modal model from the Ant Ling team.

LLaDA2.1-flash09 Feb 2026
100BApache-2.058.8K · 30d

An update to the diffusion model series by Ant Ling, with performance being roughly on par with Qwen3 30B-A3B. However, while this model is three times the size, the diffusion-based generation makes up for it in speed.

Ling-1T09 Oct 2025
1T (50B active)MIT1.6K · 30d

Inclusion AI is waking up and starting to adopt the release cadence of its bigger brother by releasing models left, right and center. Similar to Qwen, they also started to scale up model sizes considerably, hitting the 1T threshold. They also release a reasoning version and experiment with different architectures and modalities. Keep an eye on them!

Ming-UniAudio-16B30 Sep 2025
16BApache-2.060 · 30d

A MoE audio-input model.

Ming-UniVision-16B-A3B30 Sep 2025
16BApache-2.042 · 30d

A VLM MoE by InclusionAI.

Ring-1T-preview29 Sep 2025
1TMIT

Inclusion AI is waking up and starting to adopt the release cadence of its bigger brother by releasing models left, right and center. Similar to Qwen, they also started to scale up model sizes considerably, hitting the 1T threshold. They also release a reasoning version and experiment with different architectures and modalities. Keep an eye on them!

Ling-flash-2.017 Sep 2025
100B total, 6.1B activeMIT2.2K · 30d

A 100B-A6B MoE by InclusionAI.

LLaDA-MoE-7B-A1B-Instruct10 Sep 2025
7BApache-2.020.9K · 30d

A MoE diffusion model.

Ling-mini-2.008 Sep 2025
16B (1.4B active)MIT21.6K · 30d

The Ant Group continues to fly under the radar, but this release is pretty interesting in our opinion: They train relatively small MoE models with only 16B total parameters, but they share different checkpoints during the training for 5, 10, 15 and 20 trillion tokens. Furthermore, they release patches for Megatron and TransformerEngine and guides how to train the models further. On top of all of that, they also release a reasoning version.

GroveMoE-Inst18 Aug 2025
33BApache-2.039 · 30d

Another MoE model by the Ant Group.

Ling-lite-1.5-250623 Jul 2025
16.8BMIT55 · 30d

A small-ish MoE model with 17B total, 2.8B active parameters, released under the MIT license.

Ming-Lite-Omni-1.515 Jul 2025
20.3B (3B active)MIT178 · 30d

An update to the omni Ming model.

AReaL-boba-2-8B03 Jun 2025
8BApache-2.039 · 30d

Ant Labs is also releasing its experiments with RL on smaller models.

Ming-Lite-Omni28 May 2025
19BMIT89 · 30d

Another omni model using the MoE architecture.

Ling-plus-base28 Feb 2025
290BMIT72 · 30d

A series of open MoEs by the Chinese Ant Group. The tech report details their approach to training these models, including the mixing of different GPUs (some of which are not NVIDIA).