What is the best open weight reasoning model?
Asked "What is the best open weight reasoning model?", ChatGPT, Copilot, Gemini, Google AI Mode and Perplexity named 30 distinct names across 10 answers on September 8, 2026, and 8 of them in two or more answers, and the engines did not agree; the closest to a consensus, DeepSeek-R1, was named by only 4 of 5.
8 of 30 names confirmed · named in 2 or more of 10 answers · asked September 8, 2026 · 5 engines
The best open-weight reasoning model depends on specific needs, but top contenders include DeepSeek-R1, Qwen3-235B, and Kimi K3, with each excelling in different areas such as raw frontier-grade reasoning, local deployment efficiency, or coding and agentic tasks.
- 1DeepSeek-R1named in 6 of 10 answers
- 2GLM-5.2named in 3 of 10 answers
- 3DeepSeek-V3.2-Specialenamed in 2 of 10 answers
- 4Qwen3-235Bnamed in 2 of 10 answers
- 5Qwen3-Thinkingnamed in 2 of 10 answers
- 6Llama 4 Mavericknamed in 2 of 10 answers
- 7Qwen3-235B-A22Bnamed in 2 of 10 answers
- 8Phi-4named in 2 of 10 answers
- 9DeepSeeknamed in 1 of 10 answersone answer
- 10DeepSeek V3.1 Terminusnamed in 1 of 10 answersone answer
- 11DeepSeek-V4named in 1 of 10 answersone answer
- 12Kimi K3named in 1 of 10 answersone answer
- 13Alibaba Cloud Qwen Thinkingnamed in 1 of 10 answersone answer
- 14DeepSeek R1named in 1 of 10 answersone answer
- 15DeepSeek-V4 Pronamed in 1 of 10 answersone answer
- 16DeepSeek-V4-Pronamed in 1 of 10 answersone answer
- 17GLM-5.3named in 1 of 10 answersone answer
- 18Qwennamed in 1 of 10 answersone answer
- 19DeepSeek-R1-Distill-Qwen-32Bnamed in 1 of 10 answersone answer
- 20gpt-oss-120bnamed in 1 of 10 answersone answer
- 21Moonshot AI Kiminamed in 1 of 10 answersone answer
- 22OpenAI gpt-oss-120bnamed in 1 of 10 answersone answer
- 23Qwen 3.5named in 1 of 10 answersone answer
- 24Qwen3.8 Maxnamed in 1 of 10 answersone answer
- 25DeepSeek V4named in 1 of 10 answersone answer
The full measurement
- The position each of the 5 engines gave all 30 names.
- How many of the 10 answers named each of them.
- 56 sampled observations behind this ranking, and where the engines disagree.
- Fan-out — the query each engine actually searched.
- Every citation, and the sources nobody cited.