Glyd

What runs on a 141 GB GPU, exactly

H200. Open models in their original precision, no quantization, with and without Glyd.

Speed

Measured per GPU, bf16 and Glyd in the same run: RTX 4080 SUPER, RTX A6000, A10, A100 and H100.

See the benchmarks

Runs only with Glyd

bf16 does not fit; Glyd does, bit for bit
Modelbf16 needsGlyd needsWith Glyd
Qwen3-Next 80B-A3B164.5 GB112.1 GBRoom for about 1.6M tokens of context with Glyd
Llama 4 Scout 109B220.5 GB149.0 GBRoom for about 17K tokens of context with Glyd

Runs either way

Glyd leaves room for a longer context or more users
Modelbf16 needsGlyd needsThe difference
SmolLM3 3B8.4 GB6.3 GBContext room 1.9M → 2.0M tokens with Glyd
Llama 3.2 3B9.0 GB6.9 GBContext room 1.2M → 1.3M tokens with Glyd
Qwen3 4B 250710.9 GB8.3 GBContext room 957K → 974K tokens with Glyd
Mistral 7B v0.317.2 GB12.4 GBContext room 1.0M → 1.1M tokens with Glyd
Llama 3.1 8B18.7 GB13.5 GBContext room 1.0M → 1.1M tokens with Glyd
Qwen3 8B19.2 GB14.0 GBContext room 900K → 936K tokens with Glyd
Gemma 3 12B26.9 GB18.8 GBContext room 1.9M → 2.0M tokens with Glyd
Phi-4 14B32.6 GB23.0 GBContext room 585K → 632K tokens with Glyd
R1 Distill Qwen 14B32.8 GB23.3 GBContext room 608K → 656K tokens with Glyd
Mistral Small 3.2 24B51.0 GB35.1 GBContext room 617K → 714K tokens with Glyd
Gemma 4 26B-A4B53.8 GB36.8 GBContext room 2.4M → 2.8M tokens with Glyd
Gemma 3 27B57.6 GB39.6 GBContext room 1.2M → 1.4M tokens with Glyd
Qwen3.8 27B57.7 GB39.5 GBContext room 1.4M → 1.7M tokens with Glyd
Muse Glimmer 30B61.4 GB41.8 GBContext room 6.7M → 8.2M tokens with Glyd
Qwen3 30B-A3B63.5 GB43.5 GBContext room 896K → 1.1M tokens with Glyd
Qwen3 32B69.3 GB48.3 GBContext room 319K → 399K tokens with Glyd
Llama 3.3 70B145.4 GB99.0 GBContext room 25K → 166K tokens with Glyd
Qwen2.5 72B149.7 GB102.1 GBContext room 11K → 157K tokens with Glyd

Too big for 141 GB

Even compressed
Modelbf16 needsGlyd needsWhat it takes
GLM-4.5-Air 106B224.1 GB151.2 GBNeeds 151.2 GB with Glyd: 2 H100s

Questions about 141 GB GPUs

Which models run on a 141 GB GPU only with Glyd?

Qwen3-Next 80B-A3B and Llama 4 Scout 109B. With an 8K-token context they need 164.5 to 220.5 GB in bf16, more than one 141 GB GPU holds, and 112.1 to 149 GB with Glyd, every weight exact.

Is it slower than bf16?

Not measured on a 141 GB GPU yet. On the GPUs measured, one sequence at a time takes 5 to 28% less GPU time a token than bf16.

How is this different from a 4-bit GGUF?

A 4-bit model is smaller, about 5 GB for an 8B model, but its weights are rounded and its answers change. Glyd gives you the original model.

Updated Sep 27, 2026. A model fits when its weights, an 8K-token KV cache and 1.5 GB for the runtime fit in the memory nvidia-smi reports: 143,771 MiB on a 141 GB card.