
ByteDance Seed
Another Lean4 generation model.
We got a lot of flak on social media for placing Bytedance relatively low on our tier list, with us noting that they are missing a general purpose open model release. This is their first entry in this regard, coming with respectable scores across the benchmarks (and with a permissively licensed, nicely sized base model - note that Qwen 3 did NOT release their 32B base model).
A translation model by ByteDance Seed, which continues to release interesting papers. It will be interesting to see whether they do a YOLO run and release a single, big LLM in the future.
ByteDance is also stepping up its game, releasing a lot of interesting papers and models. Just like its Chinese peers, it is released under a permissive license. This small model is useful for restoring or upscaling videos.
A multimodal model using the Mixture-of-Transformer-Experts (MoT) architecture, which uses two encoders (one for pixels, one for semantic features of images). It can also generate and edit images. Truly multimodal models are picking up steam in the open-source world!
ByteDance also started dropping a lot of open models and research. Seed-Coder is a coding model, trained on 6T tokens, which are curated by using LLMs as a quality filter. The reasoning variant is trained with GRPO. Unsurprisingly for Chinese model releases, the model is released under MIT.
Parallel to the preview of OpenAI's Operator, Bytedance has released a set of GUI models used to control a computer. The models are based on Qwen2-VL and released under Apache 2.0. Alongside the models, they also released apps for both macOS and Windows to let those models control the user's computer.