A multimodal model which can process text, vision, and audio as both inputs and outputs.
Specs
Params73B total, 3B active
LicenseMIT
Resources
Adoption · Hugging Face
RAM score
Relative Adoption Metric not scored because required parameter, download, or API metadata is not cataloged. This is not a zero score.
Hugging Face Downloads
216
last 30d
30.9K
all time
HF Likes
207
Relative Adoption Metric not scored because required parameter, download, or API metadata is not cataloged. This is not a zero score.
Related Models
More from Meituan LongCat
Similar Models
Ming-Lite-Omni-1.5
Inclusion AI15 Jul 2025multimodalimage generationaudio generation
Ming-Lite-Omni
Inclusion AI28 May 2025multimodalimage generationaudio generation
HunyuanImage-3.0-Instruct
Tencent26 Jan 2026multimodalimage generation80B params
Emu3.5-Image
BAAI31 Oct 2025multimodalimage generation34.1B params
