
MiniMax
Minimax released a new, slightly bigger model with its own sparse attention variant under a revenue-gated, non-commercial license.
Minimax also plans to IPO next month and released an update to M2 ahead of this milestone. M2.1 is a notable upgrade over the previous version, especially for UI generation.
Probably one of the surprise releases this month: Minimax, whose previous models were usually behind the claimed performance, really took a leap forward with M2, putting them squarely into the spotlight. They also speedrun the (Chinese) model release playbook, something we observed in the past and was perfected by the likes of Alibaba (Qwen), Moonshot (Kimi) and Zhipu (GLM): Build a social media presence, mainly on Twitter. This increasingly means that researchers are also active aside from the corporate / brand accounts. Release a new model with (Western) launch partners and ecosystem support on day zero, from vLLM to OpenRouter and tools like Cline. To really get it off the ground, offer free access to the API for a limited time. Offer a coding subscription which is compatible with Claude Code (or fork a CLI) while undercutting their pricing. Develop your own tooling and train the next model to work even better with it. This strategy, of course, is working: Zhipu reportedly has over 100K international API users and 3M chatbot users.
Minimax has also released their first reasoning model, based on their own hybrid attention MoE architecture. They release two models, one with 40K and one with a whopping 80K thinking budget, which is the most of any available model, both open and closed. However, this does not necessarily mean that they are the best models for reasoning tasks; in our tests, the model really tends to overthink and spend a lot of tokens to reconsider its answers.
Hailuo AI, regarded as the Chinese CharacterAI, also known for their (closed-source) video generation models, has released its first text model under a permissive license (still is custom, but simpler than Llama 3 license). Similar to DeepSeek V3, it is a big MoE model with 456B total params (46B active). While the model is not as strong as DeepSeek V3, the main selling point is its 4 Million context window, which is unmatched in the open model space. In an interview (in Chinese), the CEO expressed that open sourcing accelerates technology development and that they want to continue to contribute to the community. They are also planning to release models more focused on coding and a reasoning model in the future.