Glyd

What runs on a 80 GB GPU, exactly

H100 and A100 80 GB. Open models in their original precision, no quantization, with and without Glyd.

Speed

Measured per GPU, bf16 and Glyd in the same run: RTX 4080 SUPER, RTX A6000, A10, A100 and H100.

See the benchmarks

Runs only with Glyd

bf16 does not fit; Glyd does, bit for bit

None of the models we track.

Runs either way

Glyd leaves room for a longer context or more users
Modelbf16 needsGlyd needsThe difference
SmolLM3 3B8.4 GB6.3 GBContext room 1.1M → 1.1M tokens with Glyd
Llama 3.2 3B9.0 GB6.9 GBContext room 676K → 694K tokens with Glyd
Qwen3 4B 250710.9 GB8.3 GBContext room 515K → 532K tokens with Glyd
Mistral 7B v0.317.2 GB12.4 GBContext room 530K → 566K tokens with Glyd
Llama 3.1 8B18.7 GB13.5 GBContext room 518K → 558K tokens with Glyd
Qwen3 8B19.2 GB14.0 GBContext room 458K → 493K tokens with Glyd
Gemma 3 12B26.9 GB18.8 GBContext room 909K → 1.0M tokens with Glyd
Phi-4 14B32.6 GB23.0 GBContext room 267K → 314K tokens with Glyd
R1 Distill Qwen 14B32.8 GB23.3 GBContext room 277K → 324K tokens with Glyd
Mistral Small 3.2 24B51.0 GB35.1 GBContext room 219K → 316K tokens with Glyd
Gemma 4 26B-A4B53.8 GB36.8 GBContext room 789K → 1.2M tokens with Glyd
Gemma 3 27B57.6 GB39.6 GBContext room 355K → 574K tokens with Glyd
Qwen3.8 27B57.7 GB39.5 GBContext room 433K → 711K tokens with Glyd
Muse Glimmer 30B61.4 GB41.8 GBContext room 1.8M → 3.3M tokens with Glyd
Qwen3 30B-A3B63.5 GB43.5 GBContext room 232K → 435K tokens with Glyd
Qwen3 32B69.3 GB48.3 GBContext room 70K → 150K tokens with Glyd

Too big for 80 GB

Even compressed
Modelbf16 needsGlyd needsWhat it takes
Llama 3.3 70B145.4 GB99.0 GBNeeds 99.0 GB with Glyd: try 96 GB
Qwen2.5 72B149.7 GB102.1 GBNeeds 102.1 GB with Glyd: try 96 GB
Qwen3-Next 80B-A3B164.5 GB112.1 GBNeeds 112.1 GB with Glyd: try 141 GB
Llama 4 Scout 109B220.5 GB149.0 GBNeeds 149.0 GB with Glyd: try 141 GB

Questions about 80 GB GPUs

Which models run on a 80 GB GPU only with Glyd?

None of the models we track: every one that fits with Glyd fits in bf16 too, and Glyd leaves the difference for a longer context.

Is it slower than bf16?

It depends on how many sequences run at once. On an H100 SXM, measured with Qwen2.5-7B, GPU time a token against bf16 is 5% less at 1, 3% less at 8, 8% more at 32 and 17% more at 64 sequences a step.

How is this different from a 4-bit GGUF?

A 4-bit model is smaller, about 5 GB for an 8B model, but its weights are rounded and its answers change. Glyd gives you the original model.

Updated Sep 27, 2026. A model fits when its weights, an 8K-token KV cache and 1.5 GB for the runtime fit in the memory nvidia-smi reports: 81,559 MiB on a 80 GB card.