Qwen3.8-Max, commonly referred to as Qwen3.8-2.4T-A95B, is Alibaba's flagship large language model released August 3, 2026. With 2.4 trillion total parameters (about 95 billion active per token) in a sparse MoE architecture with hybrid attention, it is the largest Qwen model ever built — roughly 10× the total parameters of the Qwen3-235B-A22B released in April 2025.
Model Specifications
| Specification | Qwen3.8-2.4T-A95B |
|---|---|
| Architecture | Sparse MoE + hybrid attention (Qwen3.5 foundation) |
| Total parameters | 2.4T |
| Active parameters | ~95B |
| Context window | 1,000,000 tokens (983,616 input / 131,072 output) |
| Reasoning budget | Up to 262,144 thinking tokens |
| Input modalities | Text, images, video, documents |
| Output modality | Text |
| API model ID | qwen3.8-max |
Key Features
- First open-weight Max-class Qwen: Qwen3.8-Max is the first Max-class Qwen promised to be open-sourced, with weights scheduled for the week of August 10, 2026 on Hugging Face and ModelScope (companion dense model Qwen3.8-27B also promised).
- Native multimodality: Text, image, and video understanding in a single model, with a top-2 ranking on the Vision Arena leaderboard.
- Long-horizon agentic capability: A 10-day autonomous coding demo and strong terminal/coding agent scores.
- Massive context: 1M-token context window with a 262K-token reasoning budget.
Benchmark Highlights
- PaperBench: 93.0 — #1 globally (vs Claude Fable 5 at 88.8, GPT-5.6 Sol at 90.5)
- OSWorld-Verified: 86.1 — #1 (vs 85.0 and 83.2)
- IFBench: 82.8 — #1 (vs 63.5 and 72.7)
- GPQA Diamond: 92.6 (vs Fable 5's 92.6, Sol's 94.1)
- Terminal-Bench 2.1: 86.6 (up from 74.5 on Qwen3.7-Max)
- Text Arena (LMArena): 1491±8 Elo, ranked #5 overall
- CodeArena WebDev: 1668 Elo, ranked #4
Pricing
| Region | Input (per 1M tokens) | Output (per 1M tokens) |
|---|---|---|
| International (QwenCloud) | $2.00 | $6.00 |
| China Mainland (DashScope) | ¥12 | ¥36 |
Batch API usage is 50% of list pricing. At $2/$6 per million tokens, Qwen3.8-Max costs roughly one-third of Claude Fable 5's list rate ($15/M output). Thinking tokens count as output tokens.
Notes
Open weights were promised for the week of August 10, 2026, but as of August 13 they had not yet appeared on the official Qwen Hugging Face organization page. The license had not been disclosed; precedent Qwen releases shipped Apache 2.0. Alibaba shares rose 7% on release day.
Conclusion
Qwen3.8-2.4T-A95B thrusts Alibaba into the top tier of global frontier models, ranking #1 on PaperBench and OSWorld while undercutting closed competitors on price. If the open-weight promise is fulfilled, it will be the largest permissively-licensed model ever released — a milestone for the open-source AI community.
Comments
No comments yet. Be the first to comment!
Related Tools
Kimi K3
www.kimi.com
Moonshot AI's open-weight 2.8T multimodal agentic model with 1M-token context, the world's first open 3T-class model rivaling closed frontier models.
DeepSeek V4 Pro 0813
www.deepseek.com
DeepSeek's flagship 1.6T MoE model with 49B active parameters, 1M-token context, MIT open weights, and world-leading coding scores at a fraction of closed-model prices.
Claude Opus 5
www.anthropic.com/claude/opus
Anthropic's frontier Opus model with 1M-token context and effort-controlled reasoning, near-Fable-5 intelligence at Opus-4.8 pricing for agents and coding.
Related Insights
After I Connected Obsidian to OpenClaw, It Started Helping Me Make Decisions
Once Obsidian stopped being just a place to store notes and started working with OpenClaw, it began helping me organize context, connect information, and improve real decisions.
Stop Cramming AI Assistants into Chat Boxes: Clawdbot Picked the Wrong Battlefield
Clawdbot is convenient, but putting it inside Slack or Discord was the wrong design choice from day one. Chat tools are not for operating tasks, and AI isn't for chatting.
The Twilight of Low-Code Platforms: Why Claude Agent SDK Will Make Dify History
A deep dive from first principles of large language models on why Claude Agent SDK will replace Dify. Exploring why describing processes in natural language is more aligned with human primitive behavior patterns, and why this is the inevitable choice in the AI era.