Glyd

Qwen3 8B

  • Qwen
  • 8.2B parameters
  • Dense
  • Apache-2.0
  • Released in bf16
  • Measured Sep 27, 2026

Qwen/Qwen3-8B on Hugging Face

Weights in GPU memory
16.4 to 11.2 GB
32% less, every weight restored bit for bit
Smallest single GPU
24 GB to 16 GB
Weights, an 8K-token KV cache and 1.5 GB for the runtime
H100 80 GB GPUs it takes
1 to 1
Weights, an 8K-token KV cache and 1.5 GB a GPU for the runtime
MMLU, 300 questions
74.0% to 74.3%
Same weights; the products sum in another order

Which single GPU it fits

Worked out from the measured weights.

MemoryGPUsbf16Glyd
16 GBRTX 4080, RTX 5080Room for about 30K tokens of context with GlydNoFits
24 GBRTX 4090, RTX 3090, A10Room for the KV cache: about 53K tokens in bf16, 88K with GlydFitsFits
32 GBRTX 5090FitsFits
48 GBRTX A6000, L40S, RTX 6000 AdaFitsFits
80 GBH100, A100 80 GBFitsFits
96 GBRTX PRO 6000, GH200FitsFits
141 GBH200FitsFits

Against DFloat11, RTX 4080 SUPER

Tokens per second, higher is better. Both hold the weights in about 11.2 GB.

DFloat11Glyd
  • 1 sequence3.4× DFloat11
    DFloat11: 13.8
    Glyd: 47.2
  • 8 sequences3.4× DFloat11
    DFloat11: 104.7
    Glyd: 360.2
  • 32 sequences3.3× DFloat11
    DFloat11: 387.5
    Glyd: 1268.6
  • 64 sequences2.5× DFloat11
    DFloat11: 746.2
    Glyd: 1897.2

Every matrix, bit for bit

Every Linear layer's matrix packed and unpacked; H100 SXM, Sep 27, 2026.

Matrices in bf1613.9 GB
With Glyd, tiered layout9.4 GB
Change−32.1%

The 12-bit layout takes 24.7 to 24.8% off each matrix and decodes with less work. The ten-model report

Same weights, same model

Every weight decodes to exactly the bf16 value it was. The matrix products add their terms in a different order than cuBLAS does, as any two GPU kernels do, so a few close answers can flip either way.

Benchmarkbf16Glyd
MMLU, 300 questions74.0%74.3%
Perplexity, enwik820.749020.7428

H100 SXM, Sep 27, 2026. Lower perplexity is better.

Run it

git clone https://github.com/surya-koritala/Glyd
cd Glyd/gpu
python e2e.py /models/Qwen3-8B --format auto --fused \
  --merge --baseline --batch 1,8,32

Today Glyd runs from its PyTorch harness; vLLM and SGLang integration is not built yet.

Also measured

Twenty more open models measured, with quality wherever bf16 fits one GPU.