What runs on a 141 GB GPU, exactly
H200. Open models in their original precision, no quantization, with and without Glyd.
Speed
Measured per GPU, bf16 and Glyd in the same run: RTX 4080 SUPER, RTX A6000, A10, A100 and H100.
See the benchmarksRuns only with Glyd
| Model | bf16 needs | Glyd needs | With Glyd |
|---|---|---|---|
| Qwen3-Next 80B-A3B | 164.5 GB | 112.1 GB | Room for about 1.6M tokens of context with Glyd |
| Llama 4 Scout 109B | 220.5 GB | 149.0 GB | Room for about 17K tokens of context with Glyd |
Runs either way
| Model | bf16 needs | Glyd needs | The difference |
|---|---|---|---|
| SmolLM3 3B | 8.4 GB | 6.3 GB | Context room 1.9M → 2.0M tokens with Glyd |
| Llama 3.2 3B | 9.0 GB | 6.9 GB | Context room 1.2M → 1.3M tokens with Glyd |
| Qwen3 4B 2507 | 10.9 GB | 8.3 GB | Context room 957K → 974K tokens with Glyd |
| Mistral 7B v0.3 | 17.2 GB | 12.4 GB | Context room 1.0M → 1.1M tokens with Glyd |
| Llama 3.1 8B | 18.7 GB | 13.5 GB | Context room 1.0M → 1.1M tokens with Glyd |
| Qwen3 8B | 19.2 GB | 14.0 GB | Context room 900K → 936K tokens with Glyd |
| Gemma 3 12B | 26.9 GB | 18.8 GB | Context room 1.9M → 2.0M tokens with Glyd |
| Phi-4 14B | 32.6 GB | 23.0 GB | Context room 585K → 632K tokens with Glyd |
| R1 Distill Qwen 14B | 32.8 GB | 23.3 GB | Context room 608K → 656K tokens with Glyd |
| Mistral Small 3.2 24B | 51.0 GB | 35.1 GB | Context room 617K → 714K tokens with Glyd |
| Gemma 4 26B-A4B | 53.8 GB | 36.8 GB | Context room 2.4M → 2.8M tokens with Glyd |
| Gemma 3 27B | 57.6 GB | 39.6 GB | Context room 1.2M → 1.4M tokens with Glyd |
| Qwen3.8 27B | 57.7 GB | 39.5 GB | Context room 1.4M → 1.7M tokens with Glyd |
| Muse Glimmer 30B | 61.4 GB | 41.8 GB | Context room 6.7M → 8.2M tokens with Glyd |
| Qwen3 30B-A3B | 63.5 GB | 43.5 GB | Context room 896K → 1.1M tokens with Glyd |
| Qwen3 32B | 69.3 GB | 48.3 GB | Context room 319K → 399K tokens with Glyd |
| Llama 3.3 70B | 145.4 GB | 99.0 GB | Context room 25K → 166K tokens with Glyd |
| Qwen2.5 72B | 149.7 GB | 102.1 GB | Context room 11K → 157K tokens with Glyd |
Too big for 141 GB
| Model | bf16 needs | Glyd needs | What it takes |
|---|---|---|---|
| GLM-4.5-Air 106B | 224.1 GB | 151.2 GB | Needs 151.2 GB with Glyd: 2 H100s |
Questions about 141 GB GPUs
Which models run on a 141 GB GPU only with Glyd?
Qwen3-Next 80B-A3B and Llama 4 Scout 109B. With an 8K-token context they need 164.5 to 220.5 GB in bf16, more than one 141 GB GPU holds, and 112.1 to 149 GB with Glyd, every weight exact.
Is it slower than bf16?
Not measured on a 141 GB GPU yet. On the GPUs measured, one sequence at a time takes 5 to 28% less GPU time a token than bf16.
How is this different from a 4-bit GGUF?
A 4-bit model is smaller, about 5 GB for an 8B model, but its weights are rounded and its answers change. Glyd gives you the original model.