Qwen2.5-Omni-7B
MultimodalAudio generation
A multimodal model from the Qwen team that can handle text, images, audio, and video inputs while generating text or audio outputs. The model features a Time-aligned Multimodal RoPE implementation for effective positional encoding of video content with synchronized audio.
Specs
Params7B
LicenseApache-2.0
Adoption · Hugging Face
RAM @ 30d
Relative Adoption Metric: 0.07×. Below benchmark. Measured at 30 days.
Hugging Face Downloads
318.3K
last 30d
6.3M
all time
HF Likes
2K
Relative Adoption Metric contextualizes downloads against the model's size bucket.
Related Models



