Glyd

gpt-oss-20b

  • OpenAI
  • Mixture of experts, 3.6B active
  • Apache-2.0
  • Released in 4-bit
  • Queued for measurement

openai/gpt-oss-20b on Hugging Face

Weights in GPU memory
13.8 to 12.6 GB
4-bit weights counted as they are: Glyd’s GPU kernels for them are not built yet est.
Smallest single GPU
16 GB to 16 GB
Weights, an 8K-token KV cache and 1.5 GB for the runtime
H100 80 GB GPUs it takes
1 to 1
Weights, an 8K-token KV cache and 1.5 GB a GPU for the runtime

Which single GPU it fits

4-bit weights counted as they are; not measured yet.

MemoryGPUsReleasedGlyd
16 GBRTX 4080, RTX 5080Room for the KV cache: about 73K tokens as released, 121K with GlydFitsFits
24 GBRTX 4090, RTX 3090, A10FitsFits
32 GBRTX 5090FitsFits
48 GBRTX A6000, L40S, RTX 6000 AdaFitsFits
80 GBH100, A100 80 GBFitsFits
96 GBRTX PRO 6000, GH200FitsFits
141 GBH200FitsFits

Measured so far

21 open models measured, with quality wherever bf16 fits one GPU.