What runs on a 80 GB GPU, exactly
H100 and A100 80 GB. Open models in their original precision, no quantization, with and without Glyd.
Speed
Measured per GPU, bf16 and Glyd in the same run: RTX 4080 SUPER, RTX A6000, A10, A100 and H100.
See the benchmarksRuns only with Glyd
Runs either way
| Model | bf16 needs | Glyd needs | The difference |
|---|---|---|---|
| SmolLM3 3B | 8.4 GB | 6.3 GB | Context room 1.1M → 1.1M tokens with Glyd |
| Llama 3.2 3B | 9.0 GB | 6.9 GB | Context room 676K → 694K tokens with Glyd |
| Qwen3 4B 2507 | 10.9 GB | 8.3 GB | Context room 515K → 532K tokens with Glyd |
| Mistral 7B v0.3 | 17.2 GB | 12.4 GB | Context room 530K → 566K tokens with Glyd |
| Llama 3.1 8B | 18.7 GB | 13.5 GB | Context room 518K → 558K tokens with Glyd |
| Qwen3 8B | 19.2 GB | 14.0 GB | Context room 458K → 493K tokens with Glyd |
| Gemma 3 12B | 26.9 GB | 18.8 GB | Context room 909K → 1.0M tokens with Glyd |
| Phi-4 14B | 32.6 GB | 23.0 GB | Context room 267K → 314K tokens with Glyd |
| R1 Distill Qwen 14B | 32.8 GB | 23.3 GB | Context room 277K → 324K tokens with Glyd |
| Mistral Small 3.2 24B | 51.0 GB | 35.1 GB | Context room 219K → 316K tokens with Glyd |
| Gemma 4 26B-A4B | 53.8 GB | 36.8 GB | Context room 789K → 1.2M tokens with Glyd |
| Gemma 3 27B | 57.6 GB | 39.6 GB | Context room 355K → 574K tokens with Glyd |
| Qwen3.8 27B | 57.7 GB | 39.5 GB | Context room 433K → 711K tokens with Glyd |
| Muse Glimmer 30B | 61.4 GB | 41.8 GB | Context room 1.8M → 3.3M tokens with Glyd |
| Qwen3 30B-A3B | 63.5 GB | 43.5 GB | Context room 232K → 435K tokens with Glyd |
| Qwen3 32B | 69.3 GB | 48.3 GB | Context room 70K → 150K tokens with Glyd |
Too big for 80 GB
| Model | bf16 needs | Glyd needs | What it takes |
|---|---|---|---|
| Llama 3.3 70B | 145.4 GB | 99.0 GB | Needs 99.0 GB with Glyd: try 96 GB |
| Qwen2.5 72B | 149.7 GB | 102.1 GB | Needs 102.1 GB with Glyd: try 96 GB |
| Qwen3-Next 80B-A3B | 164.5 GB | 112.1 GB | Needs 112.1 GB with Glyd: try 141 GB |
| Llama 4 Scout 109B | 220.5 GB | 149.0 GB | Needs 149.0 GB with Glyd: try 141 GB |
Questions about 80 GB GPUs
Which models run on a 80 GB GPU only with Glyd?
None of the models we track: every one that fits with Glyd fits in bf16 too, and Glyd leaves the difference for a longer context.
Is it slower than bf16?
It depends on how many sequences run at once. On an H100 SXM, measured with Qwen2.5-7B, GPU time a token against bf16 is 5% less at 1, 3% less at 8, 8% more at 32 and 17% more at 64 sequences a step.
How is this different from a 4-bit GGUF?
A 4-bit model is smaller, about 5 GB for an 8B model, but its weights are rounded and its answers change. Glyd gives you the original model.