HologramHologramModels Hub
Index Sep 17, 2026
Models

ukisai

Swift-Qwen3.8-27b

VerifiedNew27.8B256KVisionSafetensors
Address
Identical bytes on
Status
Verified
Trending
#10
Downloads, 30 days
3.2k
Weights
55.6 GB
Sources
2
Revision
Manifest

Swift-Qwen3.8-27B is UkisAI's reasoning-efficient derivative of Qwen3.8-27B, using 58.3% fewer thinking tokens while maintaining near-identical performance (<1% loss) and as a result getting a x1.95 speed-up on several tasks.

At a glance

Task
Vision language
Input
image, text
Output
text
Parameters
27.8B
Architecture
Qwen 3.5
Context
256K tokens
Precision
BF16
Format
Safetensors
Library
transformers
License
other
Base model
Fine tuned from Qwen/Qwen3.8-27B
Released
Sep 2026
Updated
Sep 2026
Likes
349
Downloads, all time
2,753

Architecture

Layers
64
Hidden size
5,120
Attention
24 heads, grouped query, 4 KV heads
Vocabulary
248,320
Positions
262,144
Tied embeddings
No
Vision encoder
qwen3_5

Family

Models built on Swift-Qwen3.8-27b.

Run it

Loads with Transformers AutoModelForMultimodalLM and AutoProcessor, pinned to the indexed revision.

from transformers import AutoModelForMultimodalLM, AutoProcessor

model = AutoModelForMultimodalLM.from_pretrained("ukisai/Swift-Qwen3.8-27b", revision="048328f4059015b63f860a453bf94834af0db683")
processor = AutoProcessor.from_pretrained("ukisai/Swift-Qwen3.8-27b", revision="048328f4059015b63f860a453bf94834af0db683")

Spaces

Used in 1 Spaces.

Read the full model card

UkisAI

[Website](https://ukisai.com)  • 
[Learn more](https://ukisai.com/products/swift)  • 
[GGUF](https://huggingface.co/ukisai/Swift-Qwen3.8-27B-GGUF)  • 
[Enterprise licensing](#license-and-access)

Swift-Qwen3.8-27B

Swift-Qwen3.8-27B is UkisAI's reasoning-efficient derivative of Qwen3.8-27B, using 58.3% fewer thinking tokens while maintaining near-identical performance (<1% loss) and as a result getting a x1.95 speed-up on several tasks.

The prompt is a sample from LiveCodeBench v6

Training approach

We built Swift by identifying reasoning-marker tokens that, in our analysis, trigger overthinking in Qwen’s reasoning rollouts. We then fine-tuned Qwen by penalizing usage of those tokens while it reasons.

Swift produces shorter reasoning traces. In our testing, we also observe fewer overthinking errors.

For maximum gains, Swift also includes a transfer component derived from BottleCap AI's ThinkingCap-Qwen3.6-27B.

Evaluation scope

All results below compare the Qwen3.8-27B BF16 base with the same base plus the Swift adapter.

Benchmarks

  Benchmark
  Score
  Mean tokens
  Median tokens




  Base
  Swift
  Base
  Swift
  Reduction
  Reduction

General reasoning

GPQA-Diamond88.38%88.28%15,0148,855↓ 41.0%↓ 58.3%

MMLU-Pro85.47%84.95%2,9801,603↓ 46.2%↓ 28.3%

C-Eval90.00%90.62%1,492804↓ 46.1%↓ 19.3%

IFBench73.53%71.80%8,0524,657↓ 42.2%↓ 50.5%

Mathematics

AIME 202698.67%94.00%22,01416,143↓ 26.7%↓ 50.2%

HMMT (Nov 2025)99.33%96.00%22,03215,189↓ 31.1%↓ 45.9%

Multimodal

ERQA67.45%66.30%4,1372,045↓ 50.6%↓ 54.6%

Agentic coding

Terminal-Bench 2.166.74%65.84%37,08627,272↓ 26.5%↓ 38.7%

LiveCodeBench v676.76%81.55%11,3748,615↓ 24.3%↓ 45.8%

How to reproduce

Serving: BF16 · vLLM 0.27.1 · Qwen3 parser · context 262,144 · thinking xhigh.

Sampling: temperature 1.0 · top_p 0.95 · top_k 20 · min_p 0 · presence_penalty 0 · repetition_penalty 1.

Benchmarks: averages over five seeds (0–4) per model; five trials per task for Terminal-Bench.

BenchmarkOutput cap

GPQA-Diamond100,000

MMLU-Pro100,000

C-Eval16,384

IFBench81,920

AIME 2026250,000

HMMT Nov 2025250,000

ERQA100,000

Terminal-Bench 2.1Agent/task limits

LiveCodeBench v632,768

Efficiency across and versus reasoning efforts

Qwen3.8's reasoning_effort setting lets users choose how much the model thinks. For Swift to be useful across these settings, it needs to reduce thinking while keeping accuracy close to the base. We therefore tested xhigh, medium, and low: thinking-token savings persist at every level.

  Reasoning effort
  Mean thinking reduction

Xhigh↓ 41.0%

Medium↓ 22.7%

Low↓ 25.8%

The efficiency also holds up against the base's own lower effort settings. On GPQA-Diamond (198 questions, 5 seeds, 990 paired calls), Swift at xhigh is compared with the base at xhigh and at medium:

  GPQA-Diamond
  Score
  Mean tokens
  Median tokens

Base · xhigh88.38%15,0146,642

Swift · xhigh88.28%8,8552,771

Base · medium84.14%4,4511,753

Swift retains the accuracy of xhigh while using about half the tokens, although it uses about double the tokens of medium.

Quantized models

Quantized deployment is the intended use for Swift: lower-memory weights paired with shorter reasoning. The INT4 evaluations below retain token savings across GPQA, IFBench, and AIME. On AIME, Swift matches or improves accuracy and reduces output-cap failures by 31–33%.

Benchmark / quantization
Base accuracy
Swift accuracy
Mean token reduction
Median token reduction

GPQA-Diamond
Mixed-precision quant W4A16 · thinking tokens88.69%88.38%↓ 32.1%↓ 50.2%

IFBench
Mixed-precision quant W4A16 · completion tokens72.58%71.25%↓ 30.1%↓ 38.0%

AIME 2026
Mixed-precision quant W4A16 · completion tokens84.00%84.00%↓ 19.0%↓ 37.5%

AIME 2026
AWQ INT4 · completion tokens82.67%84.00%↓ 22.8%↓ 34.8%

Quantized evaluation settings

Each row compares the same quantized base with and without the Swift adapter. GPQA and AIME use five seeds; IFBench uses four samples per prompt and strict scoring. Output caps: GPQA 100,000; IFBench 81,920; AIME 32,768. GPQA and IFBench use saved historical base runs. AIME uses template-default effort and counts truncated answers as incorrect. Its shorter cap makes it a separate comparison from the BF16 table.

How to use

GGUF download

The GGUF version is available for compatible llama.cpp-based runtime.

UkisAI API

Swift is served through an OpenAI-compatible API at https://ukisai.com/api/swift/v1. It is free for research purposes and needs no API key. The model id is swift.

from openai import OpenAI

client = OpenAI(base_url="https://ukisai.com/api/swift/v1", api_key="none")
response = client.chat.completions.create(
    model="swift",
    messages=[{"role": "user", "content": "Explain speculative decoding in two sentences."}],
)
print(response.choices[0].message.content)
curl https://ukisai.com/api/swift/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model": "swift", "messages": [{"role": "user", "content": "Hello, Swift."}]}'

Transformers

import torch
from transformers import AutoModelForImageTextToText, AutoProcessor

model_id = "ukisai/Swift-Qwen3.8-27b"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForImageTextToText.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto",
)

vLLM

vllm serve ukisai/Swift-Qwen3.8-27b \
  --dtype bfloat16 \
  --tensor-parallel-size 1 \
  --max-model-len 262144 \
  --reasoning-parser qwen3 \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_coder \
  --port 8000

SGLang

Alternatively, use a current SGLang build with Qwen3.8 support:

python -m sglang.launch_server \
  --model-path ukisai/Swift-Qwen3.8-27b \
  --dtype bfloat16 \
  --tp-size 1 \
  --context-length 262144 \
  --reasoning-parser qwen3 \
  --tool-call-parser qwen3_coder \
  --port 8000

Adjust tensor parallelism and context length to your GPU memory. See the base model's vLLM recipe and SGLang recipe for installation and hardware-specific settings.

Optional MTP decoding

The published weights include the base model's MTP head. To enable self-speculative decoding, append the corresponding flags to the server command above:

# vLLM
--speculative-config '{"method":"mtp","num_speculative_tokens":3}'

# SGLang
--speculative-algorithm EAGLE --speculative-num-steps 3 \
  --speculative-eagle-topk 1 --speculative-num-draft-tokens 4

License and access

Swift-Qwen3.8-27B is a derivative of Qwen3.8-27B (Copyright 2026 Alibaba Cloud, Apache License 2.0). UkisAI's contribution, the fine-tuned weights, is licensed under the Swift Open License v1.0. See NOTICE for exactly what was changed.

Personal, research, educational, evaluation, and commercial use are free for individuals and organizations with gross annual revenue, including affiliates, of up to US$1,000,000. Above that threshold, commercial use requires a separate Swift Enterprise License. Contact UkisAI for terms.

Nothing in the Swift Open License limits your rights in Qwen3.8-27B itself under Apache 2.0.

Citation

@misc{swift-qwen3.8-27b,
  title  = {Swift-Qwen3.8-27B},
  author = {UkisAI},
  year   = {2026},
  url    = {https://huggingface.co/ukisai/Swift-Qwen3.8-27b}
}

Acknowledgements

We acknowledge the NVIDIA Innovation Lab for providing access to 8× NVIDIA H100 GPUs to train Swift.

Derived on Sep 17, 2026 from Hugging Face at revision 048328f4, README.md , config.json .