A multimodal model capable of generating images. It uses both Qwen2.5-VL and Lumina-Next and is built upon open-source data.