Open models, and what Glyd saves on each
The open models people run most, their size as released and with Glyd, and the smallest single GPU that holds them. Measured where marked; the rest are queued.
Released in bf16−32 to −33%Most models you can run on one GPU: Llama, Qwen, Gemma, Mistral Small, Phi. Glyd’s full saving, measured on twelve.
Released in FP8−16 to −18%Measured on three FP8 models as the lossless floor. The GPU kernels for FP8 are next on the roadmap.
Released in 4-bitabout 0%gpt-oss, DeepSeek-V4, Kimi K3: their weights are already rounded to 4 bits, with little left to take.
Run on one GPU
| Model | Released | Format | As released | With Glyd | Size | Saved | One GPU, as released → Glyd | Status |
|---|---|---|---|---|---|---|---|---|
| SmolLM3 3BHugging Face · 3.1B | Jul 2025 | bf16 | 6.2 GB | 4.1 GB | −32.9% | 16 GB → 16 GB | Measured | |
| Llama 3.2 3BMeta · 3.2B | Sep 2024 | bf16 | 6.4 GB | 4.3 GB | −32.8% | 16 GB → 16 GB | Measured | |
| Qwen3 4B 2507Qwen · 4B | Aug 2025 | bf16 | 8.0 GB | 5.5 GB | −32.2% | 16 GB → 16 GB | Measured | |
| gpt-oss-20bOpenAI | Aug 2025 | 4-bit | 13.8 GB | 12.6 GB | −8.6% | 16 GB → 16 GB | Queued | |
| Mistral 7B v0.3Mistral AI · 7.2B | May 2024 | bf16 | 14.5 GB | 9.8 GB | −32.7% | 24 GB → 16 GB | Measured | |
| Llama 3.1 8BMeta · 8B | Jul 2024 | bf16 | 16.1 GB | 10.8 GB | −32.8% | 24 GB → 16 GB | Measured | |
| Qwen3 8BQwen · 8.2B | Apr 2025 | bf16 | 16.4 GB | 11.2 GB | −31.9% | 24 GB → 16 GB | Measured | |
| Gemma 3 12BGoogle · 12.2B | Mar 2025 | bf16 | 24.4 GB | 16.4 GB | −32.9% | 32 GB → 24 GB | Measured | |
| Phi-4 14BMicrosoft · 14.7B | Dec 2024 | bf16 | 29.3 GB | 19.7 GB | −32.9% | 32 GB → 24 GB | Measured | |
| R1 Distill Qwen 14BDeepSeek · 14.8B | Jan 2025 | bf16 | 29.5 GB | 20.1 GB | −31.9% | 32 GB → 24 GB | Measured | |
| Mistral Small 3.2 24BMistral AI · 24B | Jun 2025 | bf16 | 48.0 GB | 32.2 GB | −33.0% | 48 GB → 48 GB | Measured | |
| Gemma 4 26B-A4BGoogle · 25.8B | Mar 2026 | bf16 | 51.6 GB | 34.7 GB | −32.8% | 80 GB → 48 GB | Measured | |
| Gemma 3 27BGoogle · 27.4B | Mar 2025 | bf16 | 54.9 GB | 36.9 GB | −32.8% | 80 GB → 48 GB | Measured | |
| Qwen3.8 27BQwen · 27.8B | Aug 2026 | bf16 | 55.6 GB | 37.3 GB | −32.8% | 80 GB → 48 GB | Measured | |
| Muse Glimmer 30BMeta · 29.8B | Aug 2026 | bf16 | 59.5 GB | 40.0 GB | −32.8% | 80 GB → 48 GB | Measured | |
| Qwen3 30B-A3BQwen · 30.5B | Apr 2025 | bf16 | 61.1 GB | 41.1 GB | −32.7% | 80 GB → 48 GB | Measured | |
| gpt-oss-120bOpenAI | Aug 2025 | 4-bit | 65.3 GB | 63.8 GB | −2.2% | 80 GB → 80 GB | Queued | |
| Qwen3 32BQwen · 32.8B | Apr 2025 | bf16 | 65.5 GB | 44.5 GB | −32.1% | 80 GB → 48 GB | Measured | |
| Mistral Small 4 119BMistral AI · 119.4B | Jan 2026 | fp8 | 121 GB | 120 GB | −0.8% | 141 GB → 141 GB | Queued | |
| Llama 3.3 70BMeta · 70.6B | Nov 2024 | bf16 | 141 GB | 94.7 GB | −32.9% | 141 GB → 96 GB | Measured | |
| Qwen2.5 72BQwen · 72.7B | Sep 2024 | bf16 | 145 GB | 97.8 GB | −32.7% | 141 GB → 96 GB | Measured | |
| Qwen3-Next 80B-A3BQwen · 81.3B | Sep 2025 | bf16 | 163 GB | 110 GB | −32.2% | none → 141 GB | Measured | |
| Llama 4 Scout 109BMeta · 108.6B | Apr 2025 | bf16 | 217 GB | 146 GB | −32.9% | none → 141 GB | Measured | |
| GLM-4.5-Air 106BZ.ai · 110.5B | Jul 2025 | bf16 | 221 GB | 148 GB | −33.0% | none → none | Measured |
Frontier models
| Model | Released | Format | As released | With Glyd | Size | Saved | 8×H100 nodes, as released → Glyd | Status |
|---|---|---|---|---|---|---|---|---|
| Mistral Medium 3.5 128BMistral AI · 127.7B | Mar 2026 | fp8 | 134 GB | 130 GB | −2.9% | 1 → 1 node | Queued | |
| GLM-5.3 FlashZ.ai · 321.3B | Aug 2026 | fp8 | 328 GB | 324 GB | −1.4% | 1 → 1 node | Queued | |
| DeepSeek-V4.1 FlashDeepSeek | Sep 2026 | 4-bit | 510 GB | 509 GB | −0.3% | 1 → 1 node | Queued | |
| MiMo-V2.6 ProXiaomi | Sep 2026 | 4-bit | 566 GB | 559 GB | −1.2% | 1 → 1 node | Queued | |
| GLM-5.3 Flash (BF16 repo)Z.ai · 321.3B | Aug 2026 | bf16 | 643 GB | 433 GB | −32.7% | 1 → 1 node | Queued | |
| GLM-5.3Z.ai · 753.3B | Aug 2026 | fp8 | 756 GB | 754 GB | −0.2% | 2 → 2 nodes | Queued | |
| K2 Horizon 375BIFM · 379.2B | Sep 2026 | bf16 | 758 GB | 510 GB | −32.7% | 2 → 1 node | Queued | |
| MiniMax-M3MiniMax · 427B | Jun 2026 | bf16 | 854 GB | 575 GB | −32.7% | 2 → 1 node | Queued | |
| DeepSeek-V4 ProDeepSeek | Apr 2026 | 4-bit | 865 GB | 863 GB | −0.2% | 2 → 2 nodes | Queued | |
| Nemotron 3 Ultra 550BNVIDIA · 560.5B | Jun 2026 | bf16 | 1121 GB | 754 GB | −32.7% | 2 → 2 nodes | Queued | |
| GLM-5.3 (BF16 repo)Z.ai · 753.3B | Aug 2026 | bf16 | 1507 GB | 1014 GB | −32.7% | 3 → 2 nodes | Queued | |
| Kimi K3Moonshot AI | Jun 2026 | 4-bit | 1561 GB | 1524 GB | −2.4% | 3 → 3 nodes | Queued | |
| InklingThinking Machines · 952.4B | Jul 2026 | bf16 | 1905 GB | 1282 GB | −32.7% | 3 → 2 nodes | Queued |
Sizes are the checkpoints on Hugging Face. Queued rows are worked out from the parameter counts at the measured 32.7% for bf16; FP8 and 4-bit weights are counted as they are, since Glyd's GPU kernels for them are not built yet. One GPU means weights, an 8K-token KV cache and 1.5 GB for the runtime.