Glyd

What runs on a 96 GB GPU, exactly

RTX PRO 6000 and GH200. Open models in their original precision, no quantization, with and without Glyd.

Speed

Measured per GPU, bf16 and Glyd in the same run: RTX 4080 SUPER, RTX A6000, A10, A100 and H100.

See the benchmarks

Runs only with Glyd

bf16 does not fit; Glyd does, bit for bit
Modelbf16 needsGlyd needsWith Glyd
Llama 3.3 70B145.4 GB99.0 GBRoom for about 19K tokens of context with Glyd
Qwen2.5 72B149.7 GB102.1 GBRoom for about 10K tokens of context with Glyd

Runs either way

Glyd leaves room for a longer context or more users
Modelbf16 needsGlyd needsThe difference
SmolLM3 3B8.4 GB6.3 GBContext room 1.3M → 1.3M tokens with Glyd
Llama 3.2 3B9.0 GB6.9 GBContext room 825K → 843K tokens with Glyd
Qwen3 4B 250710.9 GB8.3 GBContext room 631K → 648K tokens with Glyd
Mistral 7B v0.317.2 GB12.4 GBContext room 660K → 696K tokens with Glyd
Llama 3.1 8B18.7 GB13.5 GBContext room 648K → 688K tokens with Glyd
Qwen3 8B19.2 GB14.0 GBContext room 574K → 610K tokens with Glyd
Gemma 3 12B26.9 GB18.8 GBContext room 1.2M → 1.3M tokens with Glyd
Phi-4 14B32.6 GB23.0 GBContext room 350K → 397K tokens with Glyd
R1 Distill Qwen 14B32.8 GB23.3 GBContext room 364K → 412K tokens with Glyd
Mistral Small 3.2 24B51.0 GB35.1 GBContext room 324K → 420K tokens with Glyd
Gemma 4 26B-A4B53.8 GB36.8 GBContext room 1.2M → 1.6M tokens with Glyd
Gemma 3 27B57.6 GB39.6 GBContext room 564K → 783K tokens with Glyd
Qwen3.8 27B57.7 GB39.5 GBContext room 694K → 972K tokens with Glyd
Muse Glimmer 30B61.4 GB41.8 GBContext room 3.1M → 4.6M tokens with Glyd
Qwen3 30B-A3B63.5 GB43.5 GBContext room 407K → 610K tokens with Glyd
Qwen3 32B69.3 GB48.3 GBContext room 136K → 216K tokens with Glyd

Too big for 96 GB

Even compressed
Modelbf16 needsGlyd needsWhat it takes
Qwen3-Next 80B-A3B164.5 GB112.1 GBNeeds 112.1 GB with Glyd: try 141 GB
Llama 4 Scout 109B220.5 GB149.0 GBNeeds 149.0 GB with Glyd: try 141 GB
GLM-4.5-Air 106B224.1 GB151.2 GBNeeds 151.2 GB with Glyd: 2 H100s

Questions about 96 GB GPUs

Which models run on a 96 GB GPU only with Glyd?

Llama 3.3 70B and Qwen2.5 72B. With an 8K-token context they need 145.4 to 149.7 GB in bf16, more than one 96 GB GPU holds, and 99 to 102.1 GB with Glyd, every weight exact.

Is it slower than bf16?

Not measured on a 96 GB GPU yet. On the GPUs measured, one sequence at a time takes 5 to 28% less GPU time a token than bf16.

How is this different from a 4-bit GGUF?

A 4-bit model is smaller, about 5 GB for an 8B model, but its weights are rounded and its answers change. Glyd gives you the original model.

Updated Sep 27, 2026. A model fits when its weights, an 8K-token KV cache and 1.5 GB for the runtime fit in the memory nvidia-smi reports: 97,887 MiB on a 96 GB card.