Glyd

What runs on a 48 GB GPU, exactly

RTX A6000, L40S and RTX 6000 Ada. Open models in their original precision, no quantization, with and without Glyd.

Speed

Measured per GPU, bf16 and Glyd in the same run: RTX 4080 SUPER, RTX A6000, A10, A100 and H100.

See the benchmarks

Runs only with Glyd

bf16 does not fit; Glyd does, bit for bit
Modelbf16 needsGlyd needsWith Glyd
Gemma 4 26B-A4B53.8 GB36.8 GBRoom for about 372K tokens of context with Glyd
Gemma 3 27B57.6 GB39.6 GBRoom for about 159K tokens of context with Glyd
Qwen3.8 27B57.7 GB39.5 GBRoom for about 192K tokens of context with Glyd
Muse Glimmer 30B61.4 GB41.8 GBRoom for about 743K tokens of context with Glyd
Qwen3 30B-A3B63.5 GB43.5 GBRoom for about 90K tokens of context with Glyd
Qwen3 32B69.3 GB48.3 GBRoom for about 21K tokens of context with Glyd

Runs either way

Glyd leaves room for a longer context or more users
Modelbf16 needsGlyd needsThe difference
SmolLM3 3B8.4 GB6.3 GBContext room 594K → 621K tokens with Glyd
Llama 3.2 3B9.0 GB6.9 GBContext room 379K → 398K tokens with Glyd
Qwen3 4B 250710.9 GB8.3 GBContext room 284K → 302K tokens with Glyd
Mistral 7B v0.317.2 GB12.4 GBContext room 270K → 306K tokens with Glyd
Llama 3.1 8B18.7 GB13.5 GBContext room 258K → 299K tokens with Glyd
Qwen3 8B19.2 GB14.0 GBContext room 227K → 263K tokens with Glyd
Gemma 3 12B26.9 GB18.8 GBContext room 390K → 512K tokens with Glyd
Phi-4 14B32.6 GB23.0 GBContext room 101K → 148K tokens with Glyd
R1 Distill Qwen 14B32.8 GB23.3 GBContext room 104K → 152K tokens with Glyd
Mistral Small 3.2 24B51.0 GB35.1 GBContext room 12K → 108K tokens with Glyd

Too big for 48 GB

Even compressed
Modelbf16 needsGlyd needsWhat it takes
Llama 3.3 70B145.4 GB99.0 GBNeeds 99.0 GB with Glyd: try 96 GB
Qwen2.5 72B149.7 GB102.1 GBNeeds 102.1 GB with Glyd: try 96 GB
Qwen3-Next 80B-A3B164.5 GB112.1 GBNeeds 112.1 GB with Glyd: try 141 GB
Llama 4 Scout 109B220.5 GB149.0 GBNeeds 149.0 GB with Glyd: try 141 GB

Questions about 48 GB GPUs

Which models run on a 48 GB GPU only with Glyd?

Gemma 4 26B-A4B, Gemma 3 27B, Qwen3.8 27B, Muse Glimmer 30B, Qwen3 30B-A3B and Qwen3 32B. With an 8K-token context they need 53.8 to 69.3 GB in bf16, more than one 48 GB GPU holds, and 36.8 to 48.3 GB with Glyd, every weight exact.

Is it slower than bf16?

On RTX A6000s, Qwen3-32B on one GPU with Glyd generates 24%, 28% and 5% more tokens a second than bf16 across two, at 1, 8 and 32 sequences.

How is this different from a 4-bit GGUF?

A 4-bit model is smaller, about 5 GB for an 8B model, but its weights are rounded and its answers change. Glyd gives you the original model.

Updated Sep 27, 2026. A model fits when its weights, an 8K-token KV cache and 1.5 GB for the runtime fit in the memory nvidia-smi reports: 49,140 MiB on a 48 GB card.