What runs on a 24 GB GPU, exactly
RTX 4090, RTX 3090 and A10. Open models in their original precision, no quantization, with and without Glyd.
Speed
Measured per GPU, bf16 and Glyd in the same run: RTX 4080 SUPER, RTX A6000, A10, A100 and H100.
See the benchmarksRuns only with Glyd
| Model | bf16 needs | Glyd needs | With Glyd |
|---|---|---|---|
| Gemma 3 12B | 26.9 GB | 18.8 GB | Room for about 119K tokens of context with Glyd |
| Phi-4 14B | 32.6 GB | 23.0 GB | Room for about 22K tokens of context with Glyd |
| R1 Distill Qwen 14B | 32.8 GB | 23.3 GB | Room for about 20K tokens of context with Glyd |
Runs either way
| Model | bf16 needs | Glyd needs | The difference |
|---|---|---|---|
| SmolLM3 3B | 8.4 GB | 6.3 GB | Context room 244K → 271K tokens with Glyd |
| Llama 3.2 3B | 9.0 GB | 6.9 GB | Context room 154K → 173K tokens with Glyd |
| Qwen3 4B 2507 | 10.9 GB | 8.3 GB | Context room 109K → 127K tokens with Glyd |
| Mistral 7B v0.3 | 17.2 GB | 12.4 GB | Context room 74K → 110K tokens with Glyd |
| Llama 3.1 8B | 18.7 GB | 13.5 GB | Context room 62K → 102K tokens with Glyd |
| Qwen3 8B | 19.2 GB | 14.0 GB | Context room 53K → 88K tokens with Glyd |
Too big for 24 GB
| Model | bf16 needs | Glyd needs | What it takes |
|---|---|---|---|
| Mistral Small 3.2 24B | 51.0 GB | 35.1 GB | Needs 35.1 GB with Glyd: try 48 GB |
| Gemma 4 26B-A4B | 53.8 GB | 36.8 GB | Needs 36.8 GB with Glyd: try 48 GB |
| Gemma 3 27B | 57.6 GB | 39.6 GB | Needs 39.6 GB with Glyd: try 48 GB |
| Qwen3.8 27B | 57.7 GB | 39.5 GB | Needs 39.5 GB with Glyd: try 48 GB |
Questions about 24 GB GPUs
Which models run on a 24 GB GPU only with Glyd?
Gemma 3 12B, Phi-4 14B and R1 Distill Qwen 14B. With an 8K-token context they need 26.9 to 32.8 GB in bf16, more than one 24 GB GPU holds, and 18.8 to 23.3 GB with Glyd, every weight exact.
Is it slower than bf16?
On an A10, measured, it is faster: 13 to 28% less GPU time a token for Qwen2.5-7B, from 1 to 64 sequences a step.
How is this different from a 4-bit GGUF?
A 4-bit model is smaller, about 5 GB for an 8B model, but its weights are rounded and its answers change. Glyd gives you the original model.