All models

OmniParser-v2.0

byMicrosoftMicrosoft· 12 Feb 2025
Multimodal

A model to generate bounding boxes for UI screenshots. Not only can it detect the position of icons, buttons, texts etc., but it can also parse the content and interpret whether the element is interactive, which is crucial for GUI agents.

Specs
LicenseMixed (AGPL + MIT)
Adoption · Hugging Face
RAM score
Hugging Face Downloads
1.6K
last 30d
155.8K
all time
HF Likes
1.3K

Relative Adoption Metric contextualizes downloads against the model's size bucket.

Related Models