Minimax is not the only model breaching a context size of a million tokens. They use Dual Chunk Attention with eight phases of long-context training. The accompanying blog (and technical report) goes into more detail. They also contributed to vLLM to make the models easily runnable for others, which will help more long-context tooling be built around open LMs in the future! The long-context recipe is below.
Explore other models with behavioral similarity to Qwen2.5-7B-Instruct-1M.
Relative Adoption Metric contextualizes downloads against the model's size bucket.
OpenRouter publishes daily token totals for its ~50 most-served models. Days missing from the chart mean Qwen2.5-7B-Instruct-1M fell below that cutoff — not zero usage. OpenRouter logs usage separately per dated model version and per variant (like ":free"); the hub sums them into one daily total per model.




