gpt-oss-120b
openai/gpt-oss-120b on Hugging Face
Weights in GPU memory
65.3 to 63.8 GB
4-bit weights counted as they are: Glyd’s GPU kernels for them are not built yet est.Smallest single GPU
80 GB to 80 GB
Weights, an 8K-token KV cache and 1.5 GB for the runtimeH100 80 GB GPUs it takes
1 to 1
Weights, an 8K-token KV cache and 1.5 GB a GPU for the runtimeWhich single GPU it fits
4-bit weights counted as they are; not measured yet.
| Memory | GPUs | Released | Glyd |
|---|---|---|---|
| 16 GB | RTX 4080, RTX 5080 | No | No |
| 24 GB | RTX 4090, RTX 3090, A10 | No | No |
| 32 GB | RTX 5090 | No | No |
| 48 GB | RTX A6000, L40S, RTX 6000 Ada | No | No |
| 80 GB | H100, A100 80 GBRoom for the KV cache: about 506K tokens as released, 545K with Glyd | Fits | Fits |
| 96 GB | RTX PRO 6000, GH200 | Fits | Fits |
| 141 GB | H200 | Fits | Fits |
Measured so far
21 open models measured, with quality wherever bf16 fits one GPU.