Glyd

Qwen3-Next 80B-A3B

  • Qwen
  • 81.3B parameters
  • Mixture of experts, 3B active
  • Apache-2.0
  • Released in bf16
  • Measured Sep 27, 2026

Qwen/Qwen3-Next-80B-A3B-Instruct on Hugging Face

Weights in GPU memory
163 to 110 GB
32% less, every weight restored bit for bit
Smallest single GPU
none to 141 GB
Weights, an 8K-token KV cache and 1.5 GB for the runtime
H100 80 GB GPUs it takes
2 to 2
Weights, an 8K-token KV cache and 1.5 GB a GPU for the runtime
Its matrices, packed
161.3 to 109.4 GB
−32.2%, every matrix unpacked bit for bit

Which single GPU it fits

Worked out from the measured weights.

MemoryGPUsbf16Glyd
16 GBRTX 4080, RTX 5080NoNo
24 GBRTX 4090, RTX 3090, A10NoNo
32 GBRTX 5090NoNo
48 GBRTX A6000, L40S, RTX 6000 AdaNoNo
80 GBH100, A100 80 GBNoNo
96 GBRTX PRO 6000, GH200NoNo
141 GBH200Room for about 1.6M tokens of context with GlydNoFits

Every matrix, bit for bit

Every Linear layer's matrix packed and unpacked; A10, Sep 27, 2026.

Matrices in bf16161.3 GB
With Glyd, tiered layout109.4 GB
Change−32.2%

The 12-bit layout takes 24.7 to 24.8% off each matrix and decodes with less work. The ten-model report

Run it

git clone https://github.com/surya-koritala/Glyd
cd Glyd/gpu
python e2e.py /models/Qwen3-Next-80B-A3B-Instruct --format auto --fused \
  --merge --baseline --batch 1,8,32

Today Glyd runs from its PyTorch harness; vLLM and SGLang integration is not built yet.

Also measured

Twenty more open models measured, with quality wherever bf16 fits one GPU.