Glyd

Open models, and what Glyd saves on each

The open models people run most, their size as released and with Glyd, and the smallest single GPU that holds them. Measured where marked; the rest are queued.

Updated Sep 27, 202621 measured · 16 queued
Released in bf16−32 to −33%Most models you can run on one GPU: Llama, Qwen, Gemma, Mistral Small, Phi. Glyd’s full saving, measured on twelve.
Released in FP8−16 to −18%Measured on three FP8 models as the lossless floor. The GPU kernels for FP8 are next on the roadmap.
Released in 4-bitabout 0%gpt-oss, DeepSeek-V4, Kimi K3: their weights are already rounded to 4 bits, with little left to take.
As releasedWith Glyd

Run on one GPU

ModelReleasedFormatAs releasedWith GlydSizeSavedOne GPU, as released → GlydStatus
SmolLM3 3BHugging Face · 3.1BJul 2025bf166.2 GB4.1 GB−32.9%16 GB → 16 GBMeasured
Llama 3.2 3BMeta · 3.2BSep 2024bf166.4 GB4.3 GB−32.8%16 GB → 16 GBMeasured
Qwen3 4B 2507Qwen · 4BAug 2025bf168.0 GB5.5 GB−32.2%16 GB → 16 GBMeasured
gpt-oss-20bOpenAIAug 20254-bit13.8 GB12.6 GB−8.6%16 GB → 16 GBQueued
Mistral 7B v0.3Mistral AI · 7.2BMay 2024bf1614.5 GB9.8 GB−32.7%24 GB → 16 GBMeasured
Llama 3.1 8BMeta · 8BJul 2024bf1616.1 GB10.8 GB−32.8%24 GB → 16 GBMeasured
Qwen3 8BQwen · 8.2BApr 2025bf1616.4 GB11.2 GB−31.9%24 GB → 16 GBMeasured
Gemma 3 12BGoogle · 12.2BMar 2025bf1624.4 GB16.4 GB−32.9%32 GB → 24 GBMeasured
Phi-4 14BMicrosoft · 14.7BDec 2024bf1629.3 GB19.7 GB−32.9%32 GB → 24 GBMeasured
R1 Distill Qwen 14BDeepSeek · 14.8BJan 2025bf1629.5 GB20.1 GB−31.9%32 GB → 24 GBMeasured
Mistral Small 3.2 24BMistral AI · 24BJun 2025bf1648.0 GB32.2 GB−33.0%48 GB → 48 GBMeasured
Gemma 4 26B-A4BGoogle · 25.8BMar 2026bf1651.6 GB34.7 GB−32.8%80 GB → 48 GBMeasured
Gemma 3 27BGoogle · 27.4BMar 2025bf1654.9 GB36.9 GB−32.8%80 GB → 48 GBMeasured
Qwen3.8 27BQwen · 27.8BAug 2026bf1655.6 GB37.3 GB−32.8%80 GB → 48 GBMeasured
Muse Glimmer 30BMeta · 29.8BAug 2026bf1659.5 GB40.0 GB−32.8%80 GB → 48 GBMeasured
Qwen3 30B-A3BQwen · 30.5BApr 2025bf1661.1 GB41.1 GB−32.7%80 GB → 48 GBMeasured
gpt-oss-120bOpenAIAug 20254-bit65.3 GB63.8 GB−2.2%80 GB → 80 GBQueued
Qwen3 32BQwen · 32.8BApr 2025bf1665.5 GB44.5 GB−32.1%80 GB → 48 GBMeasured
Mistral Small 4 119BMistral AI · 119.4BJan 2026fp8121 GB120 GB−0.8%141 GB → 141 GBQueued
Llama 3.3 70BMeta · 70.6BNov 2024bf16141 GB94.7 GB−32.9%141 GB → 96 GBMeasured
Qwen2.5 72BQwen · 72.7BSep 2024bf16145 GB97.8 GB−32.7%141 GB → 96 GBMeasured
Qwen3-Next 80B-A3BQwen · 81.3BSep 2025bf16163 GB110 GB−32.2%none → 141 GBMeasured
Llama 4 Scout 109BMeta · 108.6BApr 2025bf16217 GB146 GB−32.9%none → 141 GBMeasured
GLM-4.5-Air 106BZ.ai · 110.5BJul 2025bf16221 GB148 GB−33.0%none → noneMeasured

Frontier models

ModelReleasedFormatAs releasedWith GlydSizeSaved8×H100 nodes, as released → GlydStatus
Mistral Medium 3.5 128BMistral AI · 127.7BMar 2026fp8134 GB130 GB−2.9%1 → 1 nodeQueued
GLM-5.3 FlashZ.ai · 321.3BAug 2026fp8328 GB324 GB−1.4%1 → 1 nodeQueued
DeepSeek-V4.1 FlashDeepSeekSep 20264-bit510 GB509 GB−0.3%1 → 1 nodeQueued
MiMo-V2.6 ProXiaomiSep 20264-bit566 GB559 GB−1.2%1 → 1 nodeQueued
GLM-5.3 Flash (BF16 repo)Z.ai · 321.3BAug 2026bf16643 GB433 GB−32.7%1 → 1 nodeQueued
GLM-5.3Z.ai · 753.3BAug 2026fp8756 GB754 GB−0.2%2 → 2 nodesQueued
K2 Horizon 375BIFM · 379.2BSep 2026bf16758 GB510 GB−32.7%2 → 1 nodeQueued
MiniMax-M3MiniMax · 427BJun 2026bf16854 GB575 GB−32.7%2 → 1 nodeQueued
DeepSeek-V4 ProDeepSeekApr 20264-bit865 GB863 GB−0.2%2 → 2 nodesQueued
Nemotron 3 Ultra 550BNVIDIA · 560.5BJun 2026bf161121 GB754 GB−32.7%2 → 2 nodesQueued
GLM-5.3 (BF16 repo)Z.ai · 753.3BAug 2026bf161507 GB1014 GB−32.7%3 → 2 nodesQueued
Kimi K3Moonshot AIJun 20264-bit1561 GB1524 GB−2.4%3 → 3 nodesQueued
InklingThinking Machines · 952.4BJul 2026bf161905 GB1282 GB−32.7%3 → 2 nodesQueued

Sizes are the checkpoints on Hugging Face. Queued rows are worked out from the parameter counts at the measured 32.7% for bf16; FP8 and 4-bit weights are counted as they are, since Glyd's GPU kernels for them are not built yet. One GPU means weights, an 8K-token KV cache and 1.5 GB for the runtime.