The first hybrid Transformer-Mamba vision model. Similar to text-only models, they find that the combination of attention and mamba layers is superior compared to only using Mamba or attention layers. The accompanying paper goes into more detail, including ablation studies.
Specs
Params97.7M
LicenseNVIDIA Source Code License-NC
Adoption · Hugging Face
RAM score
—
Hugging Face Downloads
3.3K
last 30d
19.6K
all time
HF Likes
7
Relative Adoption Metric contextualizes downloads against the model's size bucket.
Related Models
