
StepFun
An update to Step-Flash, which is really strong in Math in particular.
StepFun really stepped up its game (no pun intended), releasing a 196B-A11B MoE with strong metrics across the board. It is especially strong in math benchmarks, beating out models that are several times larger than it.
A speech dialogue model by StepFun.
A multimodal version of Qwen3 with a rather beefy 1.7B vision encoder and 1.2T tokens of continued, multimodal training, followed by an RL stage.
An update to StepFun's image generation model.
A multi-modal MoE by StepFun, one of the Chinese AI tigers.
A big (130B) audio generation model based on Step-Omni.
The image-to-video variant of StepVideo.
Another release from StepFun: A 30B video generation model with MIT license. It is safe to say that the (open) video generation space is dominated by Chinese labs and companies.
Stepfun, a Chinese lab, has released a 132B parameter LLM, which can use audio as input and generate it as well.