What runs on a 16 GB GPU, exactly
RTX 4080, 4080 SUPER, 5080, 5070 Ti and 4060 Ti 16 GB. Open models in their original precision, no quantization, with and without Glyd.
Measured on our RTX 4080 SUPER
47.2tokens/s, Qwen3-8B, one sequence
Its 16.38 GB of bf16 leave no room to run on this card. Glyd holds it in 11.15 GB, every weight exact. DFloat11 at the same size: 13.8 tokens/s.
Runs only with Glyd
| Model | bf16 needs | Glyd needs | With Glyd |
|---|---|---|---|
| Mistral 7B v0.3 | 17.2 GB | 12.4 GB | Room for about 44K tokens of context with Glyd |
| Llama 3.1 8B | 18.7 GB | 13.5 GB | Room for about 36K tokens of context with Glyd |
| Qwen3 8B | 19.2 GB | 14.0 GB | Room for about 30K tokens of context with Glyd |
Runs either way
| Model | bf16 needs | Glyd needs | The difference |
|---|---|---|---|
| SmolLM3 3B | 8.4 GB | 6.3 GB | Context room 128K → 155K tokens with Glyd |
| Llama 3.2 3B | 9.0 GB | 6.9 GB | Context room 80K → 98K tokens with Glyd |
| Qwen3 4B 2507 | 10.9 GB | 8.3 GB | Context room 51K → 69K tokens with Glyd |
Too big for 16 GB
| Model | bf16 needs | Glyd needs | What it takes |
|---|---|---|---|
| Gemma 3 12B | 26.9 GB | 18.8 GB | Needs 18.8 GB with Glyd: try 24 GB |
| Phi-4 14B | 32.6 GB | 23.0 GB | Needs 23.0 GB with Glyd: try 24 GB |
| R1 Distill Qwen 14B | 32.8 GB | 23.3 GB | Needs 23.3 GB with Glyd: try 24 GB |
| Mistral Small 3.2 24B | 51.0 GB | 35.1 GB | Needs 35.1 GB with Glyd: try 48 GB |
Speed on this GPU
| Sequences a step | bf16 | Glyd | Against bf16 |
|---|---|---|---|
| 1 sequence | 21.97 ms | 16.52 ms | 25% less time |
| 8 sequences | 22.85 ms | 17.37 ms | 24% less time |
| 32 sequences | 26.19 ms | 18.80 ms | 28% less time |
| 64 sequences | 27.67 ms | 25.14 ms | 9% less time |
Questions about 16 GB GPUs
Can I run an 8B model on a 16 GB GPU without quantizing it?
Yes, with Glyd. Llama 3.1 8B and Qwen3 8B take about 16 GB in bf16, which leaves no room to run; Glyd holds them in about 11 GB with every weight exact.
Is it slower than bf16?
On this card it is faster: 9 to 28% less GPU time a token for Qwen2.5-7B, from 1 to 64 sequences a step.
How is this different from a 4-bit GGUF?
A 4-bit model is smaller, about 5 GB for an 8B model, but its weights are rounded and its answers change. Glyd gives you the original model.