A GUI grounding model to detect the location of objects in (computer) screens. Despite the name, it has no relation to R1 other than the training style. Their blog post goes into more detail.