Kimi K2 Thinking
The best open model, competitive with some of the best closed models. However, independent evaluation is a problem as third-party API providers struggle to implement the model correctly, something which we have seen, for example, with the release of GPT-OSS. As an example, running the agentic "Vending-Bench" from Andon Labs with a third party provider as opposed to the official API makes a huge difference:
This is a huge problem plaguing open models. Moonshot also documents the tool calling accuracy in a repo, where a lot of providers perform sub-par, including vLLM with a schema accuracy <90%.
Relative Adoption Metric contextualizes downloads against the model's size bucket.
OpenRouter publishes daily token totals for its ~50 most-served models. Days missing from the chart mean Kimi K2 Thinking fell below that cutoff — not zero usage. OpenRouter logs usage separately per dated model version and per variant (like ":free"); the hub combines provider variants while keeping separately cataloged releases distinct.



