
Mistral
A Mistral Small model fine-tuned for Lean.
A non-commercial speech-generation model.
A Lean4 fine-tune of the new Mistral Small 4.
A small speech-to-text model by Mistral, which supports 13 languages, including Chinese, English and Hindi, while matching the precision of Whisper.
A 119B-A7B model by Mistral, combining their previous model generations into one as a hybrid reasoning model with coding abilities.
A V3/R1-sized model from Mistral, trained from the ground up. The model also supports vision capabilities.
An update to Devstral by Mistral, which also, similar to Kimi and Qwen, releases an open source CLI to use the model.
Small models by Mistral. However, the legal documents reveal that Ministral 3 are pruned versions of Ministral 3.1, not entirely new models.
An update to the reasoning model by Mistral.
An updated version of Devstral - "an agentic LLM for software engineering tasks built under a collaboration between Mistral AI and All Hands AI."
The first voice models from Mistral, powered by their own LLMs. They come in 3B and 24B.
An update to the Mistral Small model.
Mistral has joined the party of reasoners with an open release building upon Mistral Small. The accompanying technical report provides a lot of details.
Mistral is back with another open model. Devstral is a fine-tuned version of Mistral Small for agentic coding tasks. These models are crucial for applications like claude-code or the open-source codex cli.
Mistral has updated their small model to support a longer context (from 32K in the last version to 128K now), as well as images as input. It keeps the Apache 2.0 license and is an overall strong model. In our testing (yes, hoping to continue to grow this internal testing), the model is stronger than a lot similarly sized models. Though, the funniest part of this release was the meme's from Mistral's patented "triangle performance plots." Below is the response from Cohere Co-founder Nick Frosst to the model, showcasing the speed of Command-A (another awesome open-weight model, highlighted below in this issue).
Mistral released a new, open model. Aside from the size inflation, they announced in their blog post that they will move away from the Mistral Research License to Apache 2.0, the license they used for their first models, like Mistral 8B and Mixtral. This is big news for opening up more downstream use!
Mistral has also updated their multimodal Pixtral model and made it available under a non-commercial license.
Nothing exceptional about Mistral's post-training, but their models are consistently solid. Small = 22.2B for Mistral. I'd like them to share more, but in the meantime, they're the default for plenty of people.
Nothing exceptional about Mistral's post-training, but their models are consistently solid. Small = 22.2B for Mistral. I'd like them to share more, but in the meantime, they're the default for plenty of people.
Mistral's first image model. Solid scores, not game-changing. What is nice is that they trained their own encoder rather than building on CLIP / openCLIP.
Mistral's frontier model that came out in the shadow of Llama 3.1. This is a very strong model, taking the same noncommercial and no-base model release approach that Cohere has been using.
Mistral collaborating with Nvidia is worth paying attention to, but nothing groundbreaking in the model.
Mistral's first RNN-based model. It's very interesting to see more labs adopting this. I can still see them in the future routing a small subset of chat queries to very specialized models.
Mistral's first math model.
Mistral is still improving their small models while we're waiting to see what their strategy update is with the likes of Llama 3 and Gemma 2 out there.