A multimodal generation model by SenseTime, one of China's first AI startups, supporting text-to-image, image-to-text, and image editing.