What runs on a 96 GB GPU, exactly
RTX PRO 6000 and GH200. Open models in their original precision, no quantization, with and without Glyd.
Speed
Measured per GPU, bf16 and Glyd in the same run: RTX 4080 SUPER, RTX A6000, A10, A100 and H100.
See the benchmarksRuns only with Glyd
| Model | bf16 needs | Glyd needs | With Glyd |
|---|---|---|---|
| Llama 3.3 70B | 145.4 GB | 99.0 GB | Room for about 19K tokens of context with Glyd |
| Qwen2.5 72B | 149.7 GB | 102.1 GB | Room for about 10K tokens of context with Glyd |
Runs either way
| Model | bf16 needs | Glyd needs | The difference |
|---|---|---|---|
| SmolLM3 3B | 8.4 GB | 6.3 GB | Context room 1.3M → 1.3M tokens with Glyd |
| Llama 3.2 3B | 9.0 GB | 6.9 GB | Context room 825K → 843K tokens with Glyd |
| Qwen3 4B 2507 | 10.9 GB | 8.3 GB | Context room 631K → 648K tokens with Glyd |
| Mistral 7B v0.3 | 17.2 GB | 12.4 GB | Context room 660K → 696K tokens with Glyd |
| Llama 3.1 8B | 18.7 GB | 13.5 GB | Context room 648K → 688K tokens with Glyd |
| Qwen3 8B | 19.2 GB | 14.0 GB | Context room 574K → 610K tokens with Glyd |
| Gemma 3 12B | 26.9 GB | 18.8 GB | Context room 1.2M → 1.3M tokens with Glyd |
| Phi-4 14B | 32.6 GB | 23.0 GB | Context room 350K → 397K tokens with Glyd |
| R1 Distill Qwen 14B | 32.8 GB | 23.3 GB | Context room 364K → 412K tokens with Glyd |
| Mistral Small 3.2 24B | 51.0 GB | 35.1 GB | Context room 324K → 420K tokens with Glyd |
| Gemma 4 26B-A4B | 53.8 GB | 36.8 GB | Context room 1.2M → 1.6M tokens with Glyd |
| Gemma 3 27B | 57.6 GB | 39.6 GB | Context room 564K → 783K tokens with Glyd |
| Qwen3.8 27B | 57.7 GB | 39.5 GB | Context room 694K → 972K tokens with Glyd |
| Muse Glimmer 30B | 61.4 GB | 41.8 GB | Context room 3.1M → 4.6M tokens with Glyd |
| Qwen3 30B-A3B | 63.5 GB | 43.5 GB | Context room 407K → 610K tokens with Glyd |
| Qwen3 32B | 69.3 GB | 48.3 GB | Context room 136K → 216K tokens with Glyd |
Too big for 96 GB
| Model | bf16 needs | Glyd needs | What it takes |
|---|---|---|---|
| Qwen3-Next 80B-A3B | 164.5 GB | 112.1 GB | Needs 112.1 GB with Glyd: try 141 GB |
| Llama 4 Scout 109B | 220.5 GB | 149.0 GB | Needs 149.0 GB with Glyd: try 141 GB |
| GLM-4.5-Air 106B | 224.1 GB | 151.2 GB | Needs 151.2 GB with Glyd: 2 H100s |
Questions about 96 GB GPUs
Which models run on a 96 GB GPU only with Glyd?
Llama 3.3 70B and Qwen2.5 72B. With an 8K-token context they need 145.4 to 149.7 GB in bf16, more than one 96 GB GPU holds, and 99 to 102.1 GB with Glyd, every weight exact.
Is it slower than bf16?
Not measured on a 96 GB GPU yet. On the GPUs measured, one sequence at a time takes 5 to 28% less GPU time a token than bf16.
How is this different from a 4-bit GGUF?
A 4-bit model is smaller, about 5 GB for an 8B model, but its weights are rounded and its answers change. Glyd gives you the original model.