Swift-Qwen3.8-27B logo

Swift-Qwen3.8-27B

Visit

Swift-Qwen3.8-27B is UkisAI's reasoning-efficient Qwen3.8-27B finetune that cuts thinking tokens about 58% with under 1% score loss.

Share:
View alternatives

Swift-Qwen3.8-27B is UkisAI's reasoning-efficient derivative of Qwen3.8-27B. The card is ukisai/Swift-Qwen3.8-27b. Hugging Face API on 2026-09-15 listed 152 likes, 459 last-month downloads, pipeline image-text-to-text, createdAt 2026-09-08, lastModified 2026-09-13. License name swift-open-license-1.0. The repo is gated. Base model is Qwen/Qwen3.8-27B (finetune). r/LocalLLaMA hot on 2026-09-15 titled it "UkisAI Swift-Qwen3.8-27B / -58.3% thinking, x1.95 speed" at 556 upvotes.

The card claim: 58.3% fewer thinking tokens, under 1% mean score loss, about x1.95 speed on several tasks. Those numbers are vendor benches. We did not rerun them.

Compare Qwen3.8-27B if you wanted the dense base, Qwen3.8-Flash-Next if you wanted a smaller local flash SKU, or Muse Glimmer if you wanted another ~30B open chat model.

Key Features

  • Shorter traces: the card says it penalizes reasoning-marker tokens that trigger overthinking, plus a transfer piece from BottleCap AI ThinkingCap-Qwen3.6-27B.
  • GPQA-Diamond (card, xhigh): base 88.38% at 15,014 mean tokens vs Swift 88.28% at 8,855. Median drop listed as 58.3%.
  • LiveCodeBench v6 (card): Swift 81.55% vs base 76.76%, with fewer tokens.
  • AIME 2026 (card): Swift 94.00% vs base 98.67%. Token cut is real; contest math lost points.
  • Serving: BF16, vLLM 0.27.1, Qwen3 parser, context 262,144, thinking xhigh. Sampling on the card: temperature 1.0, topp 0.95, topk 20.
  • Free research API: https://ukisai.com/api/swift/v1, model id swift, API key none.

Limitation: 152 likes is heat, not an audit. The card is gated. Commercial use above US$1,000,000 annual recurring revenue (affiliates included) needs a Swift Enterprise License. INT4 rows use mixed settings and a shorter AIME cap. Hardware for BF16 is a full 27B-class box.

Specs

Item Value Source
Parameters ~28B BF16 HF card, 2026-09-15
Context (serve recipe) 262,144 tokens same card
License Swift Open License v1.0, gated HF cardData
Likes / last-month downloads 152 / 459 HF API, 2026-09-15
Research API Free, no key ukisai.com
Enterprise Required above $1M ARR same card

Use Cases

  • Local Qwen3.8-27B boxes that want xhigh accuracy with fewer thinking tokens.
  • People who saw the r/LocalLLaMA 556-upvote post and need the AIME drop spelled out.
  • People who should stay on the base if they needed peak contest math, not cheaper traces.

Getting Started

  1. Accept the gate on ukisai/Swift-Qwen3.8-27b.
  2. For llama.cpp, use ukisai/Swift-Qwen3.8-27B-GGUF.
  3. Or call https://ukisai.com/api/swift/v1 with model swift for research.
  4. For vLLM, follow the card: --reasoning-parser qwen3, max len 262,144.

First-party resource: the model card and ukisai.com/products/swift.

Frequently Asked Questions

Is this Qwen3.8-27B with a LoRA left on?

The card is a finetune of Qwen3.8-27B. Quant rows compare the same quantized base with and without the Swift adapter.

Can a company use it for free?

Personal, research, education, evaluation, and commercial use are free up to US$1M ARR including affiliates. Above that, contact UkisAI.

Does Swift beat medium effort of the base?

On GPQA-Diamond the card says Swift xhigh keeps xhigh accuracy at about half the tokens, and about double the tokens of base medium (84.14%).

Alternatives

Tips

  1. Do not paste the 58.3% median as a universal speedup. It is GPQA-Diamond xhigh on the card.
  2. Keep temperature 1.0 and top_p 0.95 unless you remeasure.
  3. Pin GGUF vs BF16. Mixed W4A16 rows are a separate table.

Comments

No comments yet. Be the first to comment!