An update to the InternVL series, which uses various sized InternViT encoders and different versions of Qwen2.5 (and InternLM) for the various model sizes. In this update, they introduce multimodal pre-training, i.e., using interleaved text-image data during the pre-training stage, followed by a SFT and an RL phase. The benchmark scores of the model are very good.
Specs
Params78B
LicenseQwen
Similarity · VAIL
VAIL Fingerprint
003d:0051:006f:00b5:0104:0158:0245:44c6
Explore other models with behavioral similarity to InternVL3-78B.
Adoption · Hugging Face
RAM @ 30d
Relative Adoption Metric: 0.89×. Below benchmark. Measured at 30 days.
Hugging Face Downloads
17.5K
last 30d
6.4M
all time
HF Likes
239
Relative Adoption Metric contextualizes downloads against the model's size bucket.
Related Models



