What runs on a 48 GB GPU, exactly
RTX A6000, L40S and RTX 6000 Ada. Open models in their original precision, no quantization, with and without Glyd.
Speed
Measured per GPU, bf16 and Glyd in the same run: RTX 4080 SUPER, RTX A6000, A10, A100 and H100.
See the benchmarksRuns only with Glyd
| Model | bf16 needs | Glyd needs | With Glyd |
|---|---|---|---|
| Gemma 4 26B-A4B | 53.8 GB | 36.8 GB | Room for about 372K tokens of context with Glyd |
| Gemma 3 27B | 57.6 GB | 39.6 GB | Room for about 159K tokens of context with Glyd |
| Qwen3.8 27B | 57.7 GB | 39.5 GB | Room for about 192K tokens of context with Glyd |
| Muse Glimmer 30B | 61.4 GB | 41.8 GB | Room for about 743K tokens of context with Glyd |
| Qwen3 30B-A3B | 63.5 GB | 43.5 GB | Room for about 90K tokens of context with Glyd |
| Qwen3 32B | 69.3 GB | 48.3 GB | Room for about 21K tokens of context with Glyd |
Runs either way
| Model | bf16 needs | Glyd needs | The difference |
|---|---|---|---|
| SmolLM3 3B | 8.4 GB | 6.3 GB | Context room 594K → 621K tokens with Glyd |
| Llama 3.2 3B | 9.0 GB | 6.9 GB | Context room 379K → 398K tokens with Glyd |
| Qwen3 4B 2507 | 10.9 GB | 8.3 GB | Context room 284K → 302K tokens with Glyd |
| Mistral 7B v0.3 | 17.2 GB | 12.4 GB | Context room 270K → 306K tokens with Glyd |
| Llama 3.1 8B | 18.7 GB | 13.5 GB | Context room 258K → 299K tokens with Glyd |
| Qwen3 8B | 19.2 GB | 14.0 GB | Context room 227K → 263K tokens with Glyd |
| Gemma 3 12B | 26.9 GB | 18.8 GB | Context room 390K → 512K tokens with Glyd |
| Phi-4 14B | 32.6 GB | 23.0 GB | Context room 101K → 148K tokens with Glyd |
| R1 Distill Qwen 14B | 32.8 GB | 23.3 GB | Context room 104K → 152K tokens with Glyd |
| Mistral Small 3.2 24B | 51.0 GB | 35.1 GB | Context room 12K → 108K tokens with Glyd |
Too big for 48 GB
| Model | bf16 needs | Glyd needs | What it takes |
|---|---|---|---|
| Llama 3.3 70B | 145.4 GB | 99.0 GB | Needs 99.0 GB with Glyd: try 96 GB |
| Qwen2.5 72B | 149.7 GB | 102.1 GB | Needs 102.1 GB with Glyd: try 96 GB |
| Qwen3-Next 80B-A3B | 164.5 GB | 112.1 GB | Needs 112.1 GB with Glyd: try 141 GB |
| Llama 4 Scout 109B | 220.5 GB | 149.0 GB | Needs 149.0 GB with Glyd: try 141 GB |
Questions about 48 GB GPUs
Which models run on a 48 GB GPU only with Glyd?
Gemma 4 26B-A4B, Gemma 3 27B, Qwen3.8 27B, Muse Glimmer 30B, Qwen3 30B-A3B and Qwen3 32B. With an 8K-token context they need 53.8 to 69.3 GB in bf16, more than one 48 GB GPU holds, and 36.8 to 48.3 GB with Glyd, every weight exact.
Is it slower than bf16?
On RTX A6000s, Qwen3-32B on one GPU with Glyd generates 24%, 28% and 5% more tokens a second than bf16 across two, at 1, 8 and 32 sequences.
How is this different from a 4-bit GGUF?
A 4-bit model is smaller, about 5 GB for an 8B model, but its weights are rounded and its answers change. Glyd gives you the original model.