Glyd

What runs on a 24 GB GPU, exactly

RTX 4090, RTX 3090 and A10. Open models in their original precision, no quantization, with and without Glyd.

Speed

Measured per GPU, bf16 and Glyd in the same run: RTX 4080 SUPER, RTX A6000, A10, A100 and H100.

See the benchmarks

Runs only with Glyd

bf16 does not fit; Glyd does, bit for bit
Modelbf16 needsGlyd needsWith Glyd
Gemma 3 12B26.9 GB18.8 GBRoom for about 119K tokens of context with Glyd
Phi-4 14B32.6 GB23.0 GBRoom for about 22K tokens of context with Glyd
R1 Distill Qwen 14B32.8 GB23.3 GBRoom for about 20K tokens of context with Glyd

Runs either way

Glyd leaves room for a longer context or more users
Modelbf16 needsGlyd needsThe difference
SmolLM3 3B8.4 GB6.3 GBContext room 244K → 271K tokens with Glyd
Llama 3.2 3B9.0 GB6.9 GBContext room 154K → 173K tokens with Glyd
Qwen3 4B 250710.9 GB8.3 GBContext room 109K → 127K tokens with Glyd
Mistral 7B v0.317.2 GB12.4 GBContext room 74K → 110K tokens with Glyd
Llama 3.1 8B18.7 GB13.5 GBContext room 62K → 102K tokens with Glyd
Qwen3 8B19.2 GB14.0 GBContext room 53K → 88K tokens with Glyd

Too big for 24 GB

Even compressed
Modelbf16 needsGlyd needsWhat it takes
Mistral Small 3.2 24B51.0 GB35.1 GBNeeds 35.1 GB with Glyd: try 48 GB
Gemma 4 26B-A4B53.8 GB36.8 GBNeeds 36.8 GB with Glyd: try 48 GB
Gemma 3 27B57.6 GB39.6 GBNeeds 39.6 GB with Glyd: try 48 GB
Qwen3.8 27B57.7 GB39.5 GBNeeds 39.5 GB with Glyd: try 48 GB

Questions about 24 GB GPUs

Which models run on a 24 GB GPU only with Glyd?

Gemma 3 12B, Phi-4 14B and R1 Distill Qwen 14B. With an 8K-token context they need 26.9 to 32.8 GB in bf16, more than one 24 GB GPU holds, and 18.8 to 23.3 GB with Glyd, every weight exact.

Is it slower than bf16?

On an A10, measured, it is faster: 13 to 28% less GPU time a token for Qwen2.5-7B, from 1 to 64 sequences a step.

How is this different from a 4-bit GGUF?

A 4-bit model is smaller, about 5 GB for an 8B model, but its weights are rounded and its answers change. Glyd gives you the original model.

Updated Sep 27, 2026. A model fits when its weights, an 8K-token KV cache and 1.5 GB for the runtime fit in the memory nvidia-smi reports: 24,564 MiB on a 24 GB card.