Release · Sep 27, 2026 · Glyd
pip install "glyd[gpu]": a model 33% smaller on the GPU, every weight exact
Glyd v0.21.0 installs with pip: glyd.from_pretrained() packs a bf16 model on the GPU as it loads, every weight bit for bit; exact=True gives bf16's logits.
Glyd v0.21.0 puts its GPU kernels in the Python package: pip install "glyd[gpu]" on Linux with an NVIDIA GPU, with no clone and no compiler. glyd.from_pretrained() then loads a model released in bf16 with every Linear layer’s weights held on the GPU in 10.8 or 12.0 bits instead of 16, each one decoding to its exact bf16 value.
The code
pip install "glyd[gpu]"
import glyd
from transformers import AutoTokenizer
model = glyd.from_pretrained("Qwen/Qwen3-8B") # any bf16 checkpoint, packed on the GPU as it loads
tok = AutoTokenizer.from_pretrained("Qwen/Qwen3-8B")
out = model.generate(**tok("The history of data compression", return_tensors="pt").to(model.device), max_new_tokens=64)
print(tok.decode(out[0]))
model = glyd.from_pretrained("Qwen/Qwen3-8B", exact=True) # logits bit-identical to bf16's
glyd.save_pretrained(model, "qwen3-8b-glyd") # the packed format (glyd-v1), loads without repacking
model = glyd.from_pretrained("qwen3-8b-glyd", verify=True) # decodes every tensor, checks its sha256
print(glyd.fit("Qwen/Qwen3-32B", gpu="48GB")) # will it fit, bf16 against Glyd
It needs Linux (x86_64 or aarch64), an NVIDIA GPU from Ampere on and PyTorch 2.5 or later for CUDA 12 or 13. Every option, and the command line: getting started.
What happens when it loads
transformers reads the checkpoint as it always does, and every Linear’s weight is packed on the GPU as it arrives, in the faster layout for the GPU: on RTX 40-series cards mma, the smallest, 10.8 bits a weight, every Linear layer’s matrix 32 to 33% smaller than bf16’s (nineteen open models); on an A10, A100 or H100 mma12, the 12-bit one (which layout on which GPU). What comes back is the transformers model, and generate() runs as it does in bf16.
A mixture of experts loads through the same call. And generate(..., cache_implementation="static") compiles the model as transformers compiles bf16’s.
Lossless, and the same outputs
Lossless should mean the model you run is the model that was released. The weights are, to the bit. But the outputs also depend on how a GPU rounds each product, and Glyd’s default products round a little differently from bf16’s: their logits are not bf16’s bit for bit.
When they have to be, exact=True runs bf16’s own path: the logits are bf16’s, bit for bit, and so is every greedy token. It is slower than the default.
Neither is exactly the fp32 result, so we measured both against the same model run in fp32, at the prompt’s last position:
| Logits off fp32’s, on average | bf16 | Glyd |
|---|---|---|
| Qwen3-0.6B, RTX 4080 SUPER | 0.041 | 0.030 |
| Qwen3-1.7B, RTX 4080 SUPER | 0.032 | 0.034 |
| Qwen3-8B, RTX PRO 6000 | 0.052 | 0.051 |
The default is about as near the true values as bf16; exact=True is for when the same is the requirement.
What was measured
gpu/check_api.py ran the wheel as it ships, on an RTX 4080 SUPER with Qwen3-0.6B and Qwen3-1.7B, and on an RTX PRO 6000 Blackwell Server Edition:
exact=True: the logits bit-identical to bf16’s and 32 of 32 greedy tokens bf16’s.save_pretrained, thenfrom_pretrained(path, verify=True): 198 tensors verified, and the logits and tokens the saved model’s.python -m glyd.gpu fit Qwen/Qwen3-1.7B --gpu 16GB, by the rule the GPU pages use: bf16 needs 6.6 GB, Glyd 5.3 GB; both fit.
gpu/check_models.py makes the same checks on nine dense models, Qwen3-0.6B to Llama-3.1-8B-Instruct; all pass. generate(), one sequence unless named:
| bf16 | Glyd | |
|---|---|---|
| RTX 4080 SUPER | ||
| Qwen3-1.7B, tokens/s | 88.9 | 98.3 |
| Qwen3-1.7B compiled, tokens/s | 151.9 | 187.8 |
| Qwen3-8B compiled, tokens/s | does not fit | 55.6 |
| granite-3.1-3b-a800m, weights, GB | 6.60 | 4.61 |
| granite-3.1-3b-a800m, tokens/s | 71.1 | 90.0 |
| granite-3.1-3b-a800m, 8 prompts, tokens/s | 195.6 | 600.9 |
| RTX PRO 6000 | ||
| Qwen3-8B, tokens/s | 50.0 | 53.0 |
| Qwen3-8B compiled, tokens/s | 75.9 | 93.0 |
How Glyd’s kernels compare with DFloat11’s on the RTX 4080 SUPER: the head-to-head. Every model’s size with Glyd: the models.
What is not there yet
Update: since this release, v0.22.0 saves and loads a mixture of experts packed and packs Llama 4’s experts, v0.23.0 carries the kernels as a CUDA library, v0.24.0 runs generate() compiled by default, v0.25.0 packs and checks a model with no Python, and v0.25.1 runs long prompts on the L4 and L40S close to bf16’s time. Still to come:
- Serving through vLLM or SGLang. A vLLM plugin is next, for
vllm serve MODEL --quantization glyd(the plan). - The KV cache compression, in the harness only.
- FP8 and 4-bit checkpoints load as they are, with little to gain (why).
- Windows and macOS have no GPU package; the codec runs there (
pip install glyd).
Where Glyd is still slower than bf16: the known gaps.
The license
The GPU part, glyd.gpu and its kernels, is source-available under the Business Source License 1.1: free for personal, educational, research and other non-commercial use, commercial production use needs a license, and each version becomes Apache-2.0 four years after its release. The codec is BSD-3-Clause OR GPL-2.0, as zstd is (the licenses).