DFlash 2
DFlash 2 is Inco AI's August 18, 2026 upgrade to DFlash, a block-diffusion draft model for speculative decoding. It is not a standalone chat model. A small drafter proposes a block of tokens in one pass; the target model verifies the block. Greedy output matches the target exactly. Sampling keeps the target distribution. First-party writeup: inco.ai/blog/dflash2. Checkpoints live at huggingface.co/collections/incoai/dflash-2 and a z-lab mirror. License on the Qwen3.8-27B drafter card is Apache 2.0. Drafter size is about 2B params for Qwen3.8-27B and 3B for Muse Glimmer.
r/LocalLLaMA treated the drop as same-day news: DFlash 2 for Qwen3.8-27B and Muse Glimmer, plus user reports pairing it with vLLM and llama.cpp. That feed is why this page exists.
Key Features
- One-pass block draft: every position in the block is predicted in parallel, not token by token.
- Path selector: keeps the top 16 candidates per position and scores adjacent pairs so the block stays coherent. Inco reports about +2.0M selector params and +0.6% cycle latency versus plain DFlash on their Qwen3-4B GSM8K setup.
- Local convolution: two-tap dynamic convolutions cut suffix decay without a 15-layer backbone. Inco reports +3% params and +0.7% cycle latency for that piece.
- Lossless verify: accepted tokens are the target model's tokens.
- Engine coverage on day one: SGLang, a vLLM PR, a llama.cpp PR, and an oMLX build are listed on the blog. Confirm those PRs before you pin a production tag.
Specs
| Item | Value | Source |
|---|---|---|
| Role | Draft model only | Qwen3.8-27B-DFlash2 card |
| Qwen3.8-27B drafter | ~2B, Apache 2.0 | same card |
| Muse Glimmer drafter | ~3B | Muse-Glimmer-30B-DFlash2 |
| Mean accept length (Qwen3.8-27B) | MTP 4.28 / DSpark 3.62 / DFlash 2 4.80 | Inco table, block size 8 |
| Batch-1 throughput vs AR (Qwen3.8-27B) | 2.67x to 3.43x across GSM8K, MATH-500, HumanEval, MBPP, MT-Bench | same card, 1x H200, SGLang |
| Software price | $0 | Apache 2.0 weights |
Those speedups are vendor-measured on one H200. We did not rerun them.
Use Cases
- Local Qwen3.8-27B: you already run Qwen3.8-27B and want more tokens per verify without changing answers.
- Muse Glimmer serving: pair with Muse Glimmer instead of the official DFlash drafter Meta ships.
- Agent loops: long tool traces burn decode time. DFlash 2 is aimed at that bill, not at a chat UI.
Limitation: if your engine has no DFLASH / draft-dflash path yet, the weights sit unused. vLLM support on the blog is a PR (52816), llama.cpp is PR 27342. Treat those as moving pins.
Pricing
Weights are free. You pay GPUs and the target model. There is no Inco SaaS seat listed on the blog. For a hosted cheap decode path, compare DeepSeek V4 Flash API pricing instead of pretending DFlash 2 has a monthly plan.
Getting Started
- Read DFlash 2: Keep Drafting Parallel.
- Download
Qwen/Qwen3.8-27Bplusincoai/Qwen3.8-27B-DFlash2. - Launch SGLang with
--speculative-algorithm DFLASHand--speculative-draft-model-path incoai/Qwen3.8-27B-DFlash2. - Measure accept length on your prompts before you trust the H200 table.
Frequently Asked Questions
Can I chat with DFlash 2 alone?
No. The card says it is not a standalone language model.
Does it change Qwen3.8-27B answers?
Greedy decode is exact. Sampling is distribution-preserving per Inco.
Is this the same as DSpark?
No. DSpark is a different community drafter. Inco's Qwen3.8-27B table puts DFlash 2 ahead of both MTP and DSpark on accept length.
Alternatives
- Qwen3.8-27B native MTP: no extra drafter, lower accept length on Inco's table.
- Muse Glimmer official DFlash: Meta's first-gen drafter; DFlash 2 is the upgrade.
- DeepSeek V4 Flash: buy tokens instead of serving a 27B locally.
Tips
- Start at block size 8 (7 draft tokens) to match the published table.
- Confirm the vLLM / llama.cpp PR numbers before you script an install.
- Do not quote the 3.5 million historical DFlash download figure as a DFlash 2 metric. That number is for the whole DFlash family as of August 2026.
Conclusion
If you already serve Qwen3.8-27B or Muse Glimmer, DFlash 2 is the drafter the LocalLLaMA thread was waiting for. Install the target and the drafter together, then measure. If you wanted a chat model, open the Qwen page instead.
Comments
No comments yet. Be the first to comment!
Related Tools
Related Insights
Running low on ChatGPT Codex quota? Switch to DeepSeek or Grok inside Codex
Codex Router lets you keep your Codex workspace while using DeepSeek, Grok, and other external models. A plain-English guide to routing, quotas, and network access.
Codex on any model: magpie makes Codex Router unnecessary
magpie is a free, open-source menu bar app that runs a local gateway and puts OpenRouter, DeepSeek and your ChatGPT, Claude, Cursor, Grok and Copilot subscriptions right into Codex's own model picker, with one click and no Codex Router.

Obsidian CLI + Codex: Turn Your Vault into an Agent Knowledge Engine
Obsidian CLI gives Codex and other agents a searchable, auditable, link-aware interface to a local Vault, using real cases and reproducible workflows.