A multimodal model using the Mixture-of-Transformer-Experts (MoT) architecture, which uses two encoders (one for pixels, one for semantic features of images). It can also generate and edit images. Truly multimodal models are picking up steam in the open-source world!
Specs
Params7B
LicenseApache-2.0
Adoption · Hugging Face
RAM @ 30d
Relative Adoption Metric: 0.00×. Below benchmark. Measured at 30 days.
Hugging Face Downloads
703
last 30d
40.3K
all time
HF Likes
1.2K
Relative Adoption Metric contextualizes downloads against the model's size bucket.
Related Models




