Glyd

gpt-oss-120b

  • OpenAI
  • Mixture of experts, 5.1B active
  • Apache-2.0
  • Released in 4-bit
  • Queued for measurement

openai/gpt-oss-120b on Hugging Face

Weights in GPU memory
65.3 to 63.8 GB
4-bit weights counted as they are: Glyd’s GPU kernels for them are not built yet est.
Smallest single GPU
80 GB to 80 GB
Weights, an 8K-token KV cache and 1.5 GB for the runtime
H100 80 GB GPUs it takes
1 to 1
Weights, an 8K-token KV cache and 1.5 GB a GPU for the runtime

Which single GPU it fits

4-bit weights counted as they are; not measured yet.

MemoryGPUsReleasedGlyd
16 GBRTX 4080, RTX 5080NoNo
24 GBRTX 4090, RTX 3090, A10NoNo
32 GBRTX 5090NoNo
48 GBRTX A6000, L40S, RTX 6000 AdaNoNo
80 GBH100, A100 80 GBRoom for the KV cache: about 506K tokens as released, 545K with GlydFitsFits
96 GBRTX PRO 6000, GH200FitsFits
141 GBH200FitsFits

Measured so far

21 open models measured, with quality wherever bf16 fits one GPU.