Glyd

GLM-5.3 (BF16 repo)

  • Z.ai
  • 753.3B parameters
  • Released in bf16
  • Queued for measurement

zai-org/GLM-5.3-BF16 on Hugging Face

Weights in GPU memory
1507 to 1014 GB
About 33% less, worked out from the parameter count; run queued est.
8×H100 nodes it takes
3 to 2
Weights and an 8K-token KV cache on 80 GB H100s
H100 80 GB GPUs it takes
18 to 13
Weights, an 8K-token KV cache and 1.5 GB a GPU for the runtime

Which single GPU it fits

Worked out from the parameter count; not measured yet.

MemoryGPUsbf16Glyd
16 GBRTX 4080, RTX 5080NoNo
24 GBRTX 4090, RTX 3090, A10NoNo
32 GBRTX 5090NoNo
48 GBRTX A6000, L40S, RTX 6000 AdaNoNo
80 GBH100, A100 80 GBNoNo
96 GBRTX PRO 6000, GH200NoNo
141 GBH200NoNo

Measured so far

21 open models measured, with quality wherever bf16 fits one GPU.