
Inclusion AI
An update to the diffusion model series from Ant Group/InclusionAI. Diffusion models are exciting for local use, as personal usage usually involves small batch sizes, a setting where diffusion models really shine.
An update to the 1T MoE by Ant Ling.
Ant Ling has not forgotten how to scale and has updated its 1T model, boosting scores across the board.
A 100B-A6B multi-modal model from the Ant Ling team.
An update to the diffusion model series by Ant Ling, with performance being roughly on par with Qwen3 30B-A3B. However, while this model is three times the size, the diffusion-based generation makes up for it in speed.
Inclusion AI is waking up and starting to adopt the release cadence of its bigger brother by releasing models left, right and center. Similar to Qwen, they also started to scale up model sizes considerably, hitting the 1T threshold. They also release a reasoning version and experiment with different architectures and modalities. Keep an eye on them!
A MoE audio-input model.
A VLM MoE by InclusionAI.
Inclusion AI is waking up and starting to adopt the release cadence of its bigger brother by releasing models left, right and center. Similar to Qwen, they also started to scale up model sizes considerably, hitting the 1T threshold. They also release a reasoning version and experiment with different architectures and modalities. Keep an eye on them!
A 100B-A6B MoE by InclusionAI.
A MoE diffusion model.
The Ant Group continues to fly under the radar, but this release is pretty interesting in our opinion: They train relatively small MoE models with only 16B total parameters, but they share different checkpoints during the training for 5, 10, 15 and 20 trillion tokens. Furthermore, they release patches for Megatron and TransformerEngine and guides how to train the models further. On top of all of that, they also release a reasoning version.
Another MoE model by the Ant Group.
A small-ish MoE model with 17B total, 2.8B active parameters, released under the MIT license.
An update to the omni Ming model.
Ant Labs is also releasing its experiments with RL on smaller models.
Another omni model using the MoE architecture.
A series of open MoEs by the Chinese Ant Group. The tech report details their approach to training these models, including the mixing of different GPUs (some of which are not NVIDIA).