
Meituan LongCat
The Chinese DoorDash is back again. This time, the company released another big MoE with 1.6T parameters. While the model itself is not the most capable for its size beyond benchmarks, it was trained entirely on Ascend 910s, making it the first non-Huawei, non-toy model trained entirely on Chinese accelerators. Other Chinese chips are mostly used for inference (if at all).
A multimodal model which can process text, vision, and audio as both inputs and outputs.
A Lean4 fine-tune of the large LongCat model.
An image editing model by Meituan.
A "small" 68.5B-A3B, non-thinking MoE by the Chinese Doordash, which is a frequent guest on this series.
A version of the big LongCat model for parallel thinking, which is the key technology behind things like GPT Pro or Grok Heavy.
An image generation model by Meituan.
The Chinese Doordash also trains video generation models, why wouldn't they.
The Chinese DoorDash is a staple of artifacts at this point. They constantly release new models, new benchmarks, and new datasets. This time, they release an omni version of the LongCat model.
Yes, the Chinese Doordash continues to release models. This time, they drop the reasoning version of their MoE, coming with an interesting tech report.
Even the Chinese DoorDash is now releasing open models with in-depth tech reports.