
OpenBMB
An English and Chinese 8B model with sparse attention, supporting a 1M context window.
A text-to-speech model with voice cloning by the OpenBMB team.
An 8B hybrid reasoning model with 64K context, English and Chinese language support and sparse attention, released under Apache 2.0. These are explicitly marketed as on-device models, such as in their tech report, and fly under the radar because they're not challenging "frontier performance" like Qwen or DeepSeek.
A small vision model by the competent OpenBMB team. However, the license is rather restrictive, requiring attribution, disallowing model training and restricting commercial use (if >5,000 devices or >1M DAU) on top of usage restrictions.
A series of small-ish models trained on 8T tokens and aimed at edge deployment. They also release CUDA Kernels for those models to use them even more efficiently.
A model trained on GUIs of smartphones to control them.
MiniCPM is an omni-model by fusing together SigLIP, Whisper, ChatTTS, and Qwen, totaling 8B parameters. It is limited to Chinese and English audio outputs, but it is multilingual.
A very strong 4B model. They get an MMLU of 67! Crazy. We're not even there with 7B models yet at Ai2.
Two new late-fusion VLMs built on the Llama 3 8B backbone. Please reach out if you have experience with these