
Zyphra
A text-to-speech model by Zyphra.
A SSM-Transformer vision language model by Zyphra.
A capable MoE with a lot of architectural tricks, including its own attention variant. It was trained entirely on AMD GPUs. An 8B variant has also been released and comes with a tech report.
A capable MoE with a lot of architectural tricks, including its own attention variant. It was trained entirely on AMD GPUs. An 8B variant has also been released and comes with a tech report.
The VL variant of the ZAYA1 model.
A small coding model, building upon the R1-distilled version of Qwen2.5 1.5B by training it with PRIME.
A text-to-speech model, released under Apache 2.0, with voice cloning capabilities.
This seems to be the strongest 7B hybrid architecture language model we have to date (with open weights). Mostly, these are important to keep an eye on, but not a ton of use I know of other than by researchers and very niche domains.
Another Mamba-style model, with three architecture improvements from Zamba1 (I list these so you can see the types of changes that are made - they tend to be minor):
A hybrid model (like StripedHyena) with Mamba and Transformer blocks.