
Skywork
Another video game generation model by Skywork.
A fine-tuned version of Qwen2.5 72B by the Skywork team. This is "just another solid reasoning model," as we see so many of in 2025.
Skywork continues to be a recurring guest to the series, releasing multiple models a month. Their R1V models are their multimodal reasoning models. V3 uses InternVL3 (V2 used QwQ) and refinements to post-training.
An update to the reward models from Skywork. These are leading on most benchmarks, including the new RewardBench v2.
Another fine-tune of a Qwen model for agentic software tasks.
Another model by Skywork: A video model to "play" video games, specifically Minecraft. They also promised to open-source the dataset.
Skywork, a Singaporean startup, has been around for some time as a solid smaller contributor to the open ecosystem, but their pace of (open) model releases picked up recently. In the last episode, we featured their math reasoning model and their video generation model. They also have closed models, such as Mureka for song generation or an agent platform to rival Manus. R1V2 is one of the few open multimodal reasoning models. They use GRPO and Mixed Preference Optimization (MPO) to train the model, rivaling the performance of closed models such as Gemini 2.0 Flash or GPT-4o-mini.
A reasoning model trained in multiple stages with increasing context length, improving the performance compared to keeping the context length constant. The blog post goes into greater detail, including their entropy experiments of the actor model.
A video generation model specialized in using reference images.
A human-centric fine-tune of HunyuanVideo on films and TV clips.
The SOTA current reward model on RewardBench was trained only by filtering known preference datasets. Data is king once again. The data is Skywork-Reward-Preference-80K-v0.1.
Another MoE model.