GLM-5.3 Flash
zai-org/GLM-5.3-Flash on Hugging Face
Weights in GPU memory
328 to 324 GB
FP8 weights counted as they are: Glyd’s GPU kernels for them are not built yet est.8×H100 nodes it takes
1 to 1
Weights and an 8K-token KV cache on 80 GB H100sH100 80 GB GPUs it takes
4 to 4
Weights, an 8K-token KV cache and 1.5 GB a GPU for the runtimeWhich single GPU it fits
FP8 weights counted as they are; not measured yet.
| Memory | GPUs | Released | Glyd |
|---|---|---|---|
| 16 GB | RTX 4080, RTX 5080 | No | No |
| 24 GB | RTX 4090, RTX 3090, A10 | No | No |
| 32 GB | RTX 5090 | No | No |
| 48 GB | RTX A6000, L40S, RTX 6000 Ada | No | No |
| 80 GB | H100, A100 80 GB | No | No |
| 96 GB | RTX PRO 6000, GH200 | No | No |
| 141 GB | H200 | No | No |
Measured so far
21 open models measured, with quality wherever bf16 fits one GPU.