Glyd

Results · Sep 26, 2026 · Glyd

How much smaller Glyd makes bf16 weights, bit for bit

Glyd stores bf16 weights in about 11 bits, not 16, and returns every value exactly: every matrix of nineteen open models came out 31.9 to 33.0% smaller.

Most open models you can run on one GPU are released in bf16: 16 bits a weight. Glyd stores the same weights in about 11 bits, and gives every one of them back exactly.

What it takes off

Across nineteen popular open models every matrix came out 31.9 to 33.0% smaller (the report), Qwen3-8B and Llama-3.3-70B-Instruct among them. The smallest layout takes 10.72 to 10.89 bits a weight, 32 to 33% smaller than bf16.

Every weight decodes to exactly the bf16 value it was (how it works). Which layout runs on which GPU is on the GPU docs; to pack a whole model and check it bit for bit, see getting started.

Where it does not apply

The saving is for models released in bf16. Models released in FP8 have less to take, 17 to 18% as measured on three of them, and Glyd’s GPU kernels for FP8 are not built yet. Models released in 4-bit, such as gpt-oss, DeepSeek-V4 and Kimi-K3, have their weights already rounded to 4 bits, with little left to take. 4-bit quantization is smaller than Glyd, 4 bits a weight and their scales, but it rounds the weights and changes the model’s answers; Glyd is for when you need the model exactly as it was trained, on less hardware.