All models

Qwen2.5-Omni-7B

byQwenQwen· 22 Mar 2025
MultimodalAudio generation

A multimodal model from the Qwen team that can handle text, images, audio, and video inputs while generating text or audio outputs. The model features a Time-aligned Multimodal RoPE implementation for effective positional encoding of video content with synchronized audio.

Specs
Params7B
LicenseApache-2.0
Adoption · Hugging Face
RAM @ 30d
Relative Adoption Metric: 0.07×. Below benchmark. Measured at 30 days.
Hugging Face Downloads
314K
last 30d
5.8M
all time
HF Likes
1.9K

Relative Adoption Metric contextualizes downloads against the model's size bucket.

Related Models