Qwen3.8-Max, commonly referred to as Qwen3.8-2.4T-A95B, is Alibaba's flagship large language model released August 3, 2026. With 2.4 trillion total parameters (about 95 billion active per token) in a sparse MoE architecture with hybrid attention, it is the largest Qwen model ever built, roughly 10x the total parameters of the Qwen3-235B-A22B released in April 2025.
Confirm the exact 2.4T SKU id before you quote a third row.
Model Specifications
| Specification | Qwen3.8-2.4T-A95B |
|---|---|
| Architecture | Sparse MoE + hybrid attention (Qwen3.5 foundation) |
| Total parameters | 2.4T |
| Active parameters | ~95B |
| Context window | 1,000,000 tokens (983,616 input / 131,072 output) |
| Reasoning budget | Up to 262,144 thinking tokens |
| Input modalities | Text, images, video, documents |
| Output modality | Text |
| API model ID | qwen3.8-max |
Key Features
- First open-weight Max-class Qwen: Qwen3.8-Max is the first Max-class Qwen promised to be open-sourced, with weights scheduled for the week of August 10, 2026 on Hugging Face and ModelScope (companion dense model Qwen3.8-27B also promised).
- Native multimodality: Text, image, and video understanding in a single model, with a top-2 ranking on the Vision Arena leaderboard.
- Long-horizon agentic capability: A 10-day autonomous coding demo and strong terminal/coding agent scores.
- Massive context: 1M-token context window with a 262K-token reasoning budget.
Benchmark Highlights
- PaperBench: 93.0 - #1 globally (vs Claude Fable 5 at 88.8, GPT-5.6 Sol at 90.5)
- OSWorld-Verified: 86.1 - #1 (vs 85.0 and 83.2)
- IFBench: 82.8 - #1 (vs 63.5 and 72.7)
- GPQA Diamond: 92.6 (vs Fable 5's 92.6, Sol's 94.1)
- Terminal-Bench 2.1: 86.6 (up from 74.5 on Qwen3.7-Max)
- Text Arena (LMArena): 1491±8 Elo, ranked #5 overall
- CodeArena WebDev: 1668 Elo, ranked #4
Pricing
| Region | Input (per 1M tokens) | Output (per 1M tokens) |
|---|---|---|
| International (QwenCloud) | $2.00 | $6.00 |
| China Mainland (DashScope) | ¥12 | ¥36 |
Batch API usage is 50% of list pricing. At $2/$6 per million tokens, Qwen3.8-Max costs roughly one-third of Claude Fable 5's list rate ($15/M output). Thinking tokens count as output tokens.
Notes
Open weights were promised for the week of August 10, 2026, but as of August 13 they had not yet appeared on the official Qwen Hugging Face organization page. The license had not been disclosed; precedent Qwen releases shipped Apache 2.0. Alibaba shares rose 7% on release day.
Conclusion
Qwen3.8-2.4T-A95B thrusts Alibaba into the top tier of global frontier models, ranking #1 on PaperBench and OSWorld while undercutting closed competitors on price. If the open-weight promise is fulfilled, it will be the largest permissively-licensed model ever released - a milestone for the open-source AI community.
Related: Claude 3.5 Sonnet and Claude 3 Haiku. Hub: models.
Comments
No comments yet. Be the first to comment!
Related Tools
Related Insights

Anthropic Subagent: The Multi-Agent Architecture Revolution
Deep dive into Anthropic multi-agent architecture design. Learn how Subagents break through context window limitations, achieve 90% performance improvements, and real-world applications in Claude Code.
Stop Cramming AI Assistants into Chat Boxes: Clawdbot Picked the Wrong Battlefield
Clawdbot is convenient, but putting it inside Slack or Discord was the wrong design choice from day one. Chat tools are not for operating tasks, and AI isn't for chatting.

Grok Bot and Hermes Bot: one person finally gets a think tank and a secretariat
Grok Bot now ships with Cursor Pro+. Hermes Bot runs on a VPS. They are not smarter chat boxes. The think tank advises, the secretariat executes, and you still make the call.