What runs on a 32 GB GPU, exactly
RTX 5090. Open models in their original precision, no quantization, with and without Glyd.
Speed
Measured per GPU, bf16 and Glyd in the same run: RTX 4080 SUPER, RTX A6000, A10, A100 and H100.
See the benchmarksRuns only with Glyd
Runs either way
| Model | bf16 needs | Glyd needs | The difference |
|---|---|---|---|
| SmolLM3 3B | 8.4 GB | 6.3 GB | Context room 358K → 386K tokens with Glyd |
| Llama 3.2 3B | 9.0 GB | 6.9 GB | Context room 228K → 246K tokens with Glyd |
| Qwen3 4B 2507 | 10.9 GB | 8.3 GB | Context room 166K → 184K tokens with Glyd |
| Mistral 7B v0.3 | 17.2 GB | 12.4 GB | Context room 138K → 174K tokens with Glyd |
| Llama 3.1 8B | 18.7 GB | 13.5 GB | Context room 126K → 166K tokens with Glyd |
| Qwen3 8B | 19.2 GB | 14.0 GB | Context room 110K → 145K tokens with Glyd |
| Gemma 3 12B | 26.9 GB | 18.8 GB | Context room 125K → 248K tokens with Glyd |
| Phi-4 14B | 32.6 GB | 23.0 GB | Context room 16K → 63K tokens with Glyd |
| R1 Distill Qwen 14B | 32.8 GB | 23.3 GB | Context room 15K → 63K tokens with Glyd |
Too big for 32 GB
| Model | bf16 needs | Glyd needs | What it takes |
|---|---|---|---|
| Mistral Small 3.2 24B | 51.0 GB | 35.1 GB | Needs 35.1 GB with Glyd: try 48 GB |
| Gemma 4 26B-A4B | 53.8 GB | 36.8 GB | Needs 36.8 GB with Glyd: try 48 GB |
| Gemma 3 27B | 57.6 GB | 39.6 GB | Needs 39.6 GB with Glyd: try 48 GB |
| Qwen3.8 27B | 57.7 GB | 39.5 GB | Needs 39.5 GB with Glyd: try 48 GB |
Questions about 32 GB GPUs
Which models run on a 32 GB GPU only with Glyd?
None of the models we track: every one that fits with Glyd fits in bf16 too, and Glyd leaves the difference for a longer context.
Is it slower than bf16?
Not measured on a 32 GB GPU yet. On the GPUs measured, one sequence at a time takes 5 to 28% less GPU time a token than bf16.
How is this different from a 4-bit GGUF?
A 4-bit model is smaller, about 5 GB for an 8B model, but its weights are rounded and its answers change. Glyd gives you the original model.