GLM-4.5-Air 106B
zai-org/GLM-4.5-Air on Hugging Face
Weights in GPU memory
221 to 148 GB
33% less, every weight restored bit for bitSmallest single GPU
none to none
Weights, an 8K-token KV cache and 1.5 GB for the runtimeH100 80 GB GPUs it takes
3 to 2
Weights, an 8K-token KV cache and 1.5 GB a GPU for the runtimeIts matrices, packed
215.9 to 144.6 GB
−33.0%, every matrix unpacked bit for bitWhich single GPU it fits
Worked out from the measured weights.
| Memory | GPUs | bf16 | Glyd |
|---|---|---|---|
| 16 GB | RTX 4080, RTX 5080 | No | No |
| 24 GB | RTX 4090, RTX 3090, A10 | No | No |
| 32 GB | RTX 5090 | No | No |
| 48 GB | RTX A6000, L40S, RTX 6000 Ada | No | No |
| 80 GB | H100, A100 80 GB | No | No |
| 96 GB | RTX PRO 6000, GH200 | No | No |
| 141 GB | H200 | No | No |
Every matrix, bit for bit
Every Linear layer's matrix packed and unpacked; A10, Sep 27, 2026.
| Matrices in bf16 | 215.9 GB |
|---|---|
| With Glyd, tiered layout | 144.6 GB |
| Change | −33.0% |
The 12-bit layout takes 24.7 to 24.8% off each matrix and decodes with less work. The ten-model report
Run it
git clone https://github.com/surya-koritala/Glyd
cd Glyd/gpu
python e2e.py /models/GLM-4.5-Air --format auto --fused \
--merge --baseline --batch 1,8,32Today Glyd runs from its PyTorch harness; vLLM and SGLang integration is not built yet.
Also measured
Twenty more open models measured, with quality wherever bf16 fits one GPU.