All models

Qwen2.5-7B-Instruct-1M

byQwenQwen· 23 Jan 2025
General purpose

Minimax is not the only model breaching a context size of a million tokens. They use Dual Chunk Attention with eight phases of long-context training. The accompanying blog (and technical report) goes into more detail. They also contributed to vLLM to make the models easily runnable for others, which will help more long-context tooling be built around open LMs in the future! The long-context recipe is below.

Specs
Params7B
LicenseApache-2.0
Similarity · VAIL
VAIL
VAIL Fingerprint
003f:0059:0073:00a4:00ce:01a1:03d5:5aad

Explore other models with behavioral similarity to Qwen2.5-7B-Instruct-1M.

Adoption · Hugging Face
RAM score
Hugging Face Downloads
75.9K
last 30d
10.5M
all time
HF Likes
376

Relative Adoption Metric contextualizes downloads against the model's size bucket.

Inference · OpenRouter
Tokens/Day
below top 50 this week
Peak Tokens/Day
4.7B
06 Mar 2025
Peak Rank
#8
06 Mar 2025

OpenRouter publishes daily token totals for its ~50 most-served models. Days missing from the chart mean Qwen2.5-7B-Instruct-1M fell below that cutoff — not zero usage. OpenRouter logs usage separately per dated model version and per variant (like ":free"); the hub sums them into one daily total per model.

Related Models