Swift-Qwen3.8-27B is UkisAI's reasoning-efficient derivative of Qwen3.8-27B. The card is ukisai/Swift-Qwen3.8-27b. Hugging Face API on 2026-09-15 listed 152 likes, 459 last-month downloads, pipeline image-text-to-text, createdAt 2026-09-08, lastModified 2026-09-13. License name swift-open-license-1.0. The repo is gated. Base model is Qwen/Qwen3.8-27B (finetune). r/LocalLLaMA hot on 2026-09-15 titled it "UkisAI Swift-Qwen3.8-27B / -58.3% thinking, x1.95 speed" at 556 upvotes.
The card claim: 58.3% fewer thinking tokens, under 1% mean score loss, about x1.95 speed on several tasks. Those numbers are vendor benches. We did not rerun them.
Compare Qwen3.8-27B if you wanted the dense base, Qwen3.8-Flash-Next if you wanted a smaller local flash SKU, or Muse Glimmer if you wanted another ~30B open chat model.
Key Features
- Shorter traces: the card says it penalizes reasoning-marker tokens that trigger overthinking, plus a transfer piece from BottleCap AI ThinkingCap-Qwen3.6-27B.
- GPQA-Diamond (card, xhigh): base 88.38% at 15,014 mean tokens vs Swift 88.28% at 8,855. Median drop listed as 58.3%.
- LiveCodeBench v6 (card): Swift 81.55% vs base 76.76%, with fewer tokens.
- AIME 2026 (card): Swift 94.00% vs base 98.67%. Token cut is real; contest math lost points.
- Serving: BF16, vLLM 0.27.1, Qwen3 parser, context 262,144, thinking xhigh. Sampling on the card: temperature 1.0, topp 0.95, topk 20.
- Free research API:
https://ukisai.com/api/swift/v1, model idswift, API keynone.
Limitation: 152 likes is heat, not an audit. The card is gated. Commercial use above US$1,000,000 annual recurring revenue (affiliates included) needs a Swift Enterprise License. INT4 rows use mixed settings and a shorter AIME cap. Hardware for BF16 is a full 27B-class box.
Specs
| Item | Value | Source |
|---|---|---|
| Parameters | ~28B BF16 | HF card, 2026-09-15 |
| Context (serve recipe) | 262,144 tokens | same card |
| License | Swift Open License v1.0, gated | HF cardData |
| Likes / last-month downloads | 152 / 459 | HF API, 2026-09-15 |
| Research API | Free, no key | ukisai.com |
| Enterprise | Required above $1M ARR | same card |
Use Cases
- Local Qwen3.8-27B boxes that want xhigh accuracy with fewer thinking tokens.
- People who saw the r/LocalLLaMA 556-upvote post and need the AIME drop spelled out.
- People who should stay on the base if they needed peak contest math, not cheaper traces.
Getting Started
- Accept the gate on ukisai/Swift-Qwen3.8-27b.
- For llama.cpp, use ukisai/Swift-Qwen3.8-27B-GGUF.
- Or call
https://ukisai.com/api/swift/v1with modelswiftfor research. - For vLLM, follow the card:
--reasoning-parser qwen3, max len 262,144.
First-party resource: the model card and ukisai.com/products/swift.
Frequently Asked Questions
Is this Qwen3.8-27B with a LoRA left on?
The card is a finetune of Qwen3.8-27B. Quant rows compare the same quantized base with and without the Swift adapter.
Can a company use it for free?
Personal, research, education, evaluation, and commercial use are free up to US$1M ARR including affiliates. Above that, contact UkisAI.
Does Swift beat medium effort of the base?
On GPQA-Diamond the card says Swift xhigh keeps xhigh accuracy at about half the tokens, and about double the tokens of base medium (84.14%).
Alternatives
- Qwen3.8-27B: dense base without the Swift license gate.
- Qwen3.8-Flash-Next: smaller local flash path.
- DeepSeek-V4-Flash: hosted flash SKU if you did not want 27B weights.
Tips
- Do not paste the 58.3% median as a universal speedup. It is GPQA-Diamond xhigh on the card.
- Keep temperature 1.0 and top_p 0.95 unless you remeasure.
- Pin GGUF vs BF16. Mixed W4A16 rows are a separate table.
Comments
No comments yet. Be the first to comment!
Related Tools
Related Insights

Grok Bot and Hermes Bot: one person finally gets a think tank and a secretariat
Grok Bot now ships with Cursor Pro+. Hermes Bot runs on a VPS. They are not smarter chat boxes. The think tank advises, the secretariat executes, and you still make the call.
Hook OpenCode Go into Codex on Windows. Do Not Open a Second Toolkit.
A ChatGPT-signed Codex desktop app still shows mostly GPT in the picker. On Windows, enable only OpenCode Go and the Grok, GLM, Kimi, DeepSeek, and MiniMax models you already pay for appear in the same selector. Keys stay local. Native GPT stays put.
What Locks Codex Is the Picker, Not the Models
You already pay for OpenCode Go, Grok, and Z.ai, but the Codex picker still shows mostly GPT. The community project codex-router does not teach another install ritual. It puts subscriptions you already bought back into the selector, keeps keys on the machine, and leaves native GPT alone.