All models

Aria

byRhymes AIRhymes AI· 26 Sep 2024
General purposeMultimodal

A new architecture for multimodal models where the vision encoder is trained natively and not just late-fusion: "Aria processes text, images, video, and code all at once, without needing separate setups for each type, demonstrating the advantages of a multimodal native model." The performance is very strong, but I'm annoyed they didn't include comparisons to Molmo at launch.

Similarity · VAIL
VAIL
VAIL Fingerprint
177d:18f1:1aed:1cef:200a:2b91:303e:46a3

Explore other models with behavioral similarity to Aria.

Adoption · Hugging Face
RAM score
Relative Adoption Metric not scored because required parameter, download, or API metadata is not cataloged. This is not a zero score.
Hugging Face Downloads
97.1K
last 30d
1.1M
all time
HF Likes
638

Relative Adoption Metric not scored because required parameter, download, or API metadata is not cataloged. This is not a zero score.

Related Models