All models

MoshiVis

byKyutaiKyutai· 27 Jan 2025
MultimodalAudio generation

An image-capable extension of the text-to-speech moshi model, which uses PaliGemma 2 as its encoder.

Specs
Params7B
LicenseCC-BY-4.0
Adoption · Hugging Face
RAM score
Relative Adoption Metric not scored because required parameter, download, or API metadata is not cataloged. This is not a zero score.
Hugging Face Downloads
0
last 30d
0
all time
HF Likes
59

Relative Adoption Metric not scored because required parameter, download, or API metadata is not cataloged. This is not a zero score.

Related Models