Apple has released a suite of vision encoders under a permissive license, surpassing both CLIP and SigLIP. Instead of using both a text and a vision encoder which are then trained to maximize the similarity of the representations, Apple uses a single vision encoder, followed by a shared multimodal decoder, which processes both the image patches and the text in sequence
Specs
Params0.3B
LicenseApple AMLR
Tags
Resources
Adoption · Hugging Face
RAM score
Relative Adoption Metric not applicable.
Hugging Face Downloads
3.7K
last 30d
45.1K
all time
HF Likes
62
Relative Adoption Metric not applicable.
Related Models




