Glyd

What runs on a 32 GB GPU, exactly

RTX 5090. Open models in their original precision, no quantization, with and without Glyd.

Speed

Measured per GPU, bf16 and Glyd in the same run: RTX 4080 SUPER, RTX A6000, A10, A100 and H100.

See the benchmarks

Runs only with Glyd

bf16 does not fit; Glyd does, bit for bit

None of the models we track.

Runs either way

Glyd leaves room for a longer context or more users
Modelbf16 needsGlyd needsThe difference
SmolLM3 3B8.4 GB6.3 GBContext room 358K → 386K tokens with Glyd
Llama 3.2 3B9.0 GB6.9 GBContext room 228K → 246K tokens with Glyd
Qwen3 4B 250710.9 GB8.3 GBContext room 166K → 184K tokens with Glyd
Mistral 7B v0.317.2 GB12.4 GBContext room 138K → 174K tokens with Glyd
Llama 3.1 8B18.7 GB13.5 GBContext room 126K → 166K tokens with Glyd
Qwen3 8B19.2 GB14.0 GBContext room 110K → 145K tokens with Glyd
Gemma 3 12B26.9 GB18.8 GBContext room 125K → 248K tokens with Glyd
Phi-4 14B32.6 GB23.0 GBContext room 16K → 63K tokens with Glyd
R1 Distill Qwen 14B32.8 GB23.3 GBContext room 15K → 63K tokens with Glyd

Too big for 32 GB

Even compressed
Modelbf16 needsGlyd needsWhat it takes
Mistral Small 3.2 24B51.0 GB35.1 GBNeeds 35.1 GB with Glyd: try 48 GB
Gemma 4 26B-A4B53.8 GB36.8 GBNeeds 36.8 GB with Glyd: try 48 GB
Gemma 3 27B57.6 GB39.6 GBNeeds 39.6 GB with Glyd: try 48 GB
Qwen3.8 27B57.7 GB39.5 GBNeeds 39.5 GB with Glyd: try 48 GB

Questions about 32 GB GPUs

Which models run on a 32 GB GPU only with Glyd?

None of the models we track: every one that fits with Glyd fits in bf16 too, and Glyd leaves the difference for a longer context.

Is it slower than bf16?

Not measured on a 32 GB GPU yet. On the GPUs measured, one sequence at a time takes 5 to 28% less GPU time a token than bf16.

How is this different from a 4-bit GGUF?

A 4-bit model is smaller, about 5 GB for an 8B model, but its weights are rounded and its answers change. Glyd gives you the original model.

Updated Sep 27, 2026. A model fits when its weights, an 8K-token KV cache and 1.5 GB for the runtime fit in the memory nvidia-smi reports: 32,607 MiB on a 32 GB card.