Release · Sep 29, 2026 · Glyd
Glyd v0.25.0: glyd pack and verify in Rust, 12-bit saves, a faster 12-bit layout
Glyd v0.25.0 packs and checks a model on the CPU with glyd pack and glyd verify, loads 12-bit saves as saved, and runs the 12-bit layout faster.
Glyd v0.25.0 packs and checks a model from the command line, with no Python or GPU; saves models in the 12-bit layout and loads them as saved; and runs the 12-bit layout faster, every product bit for bit the same. Every weight still decodes to its exact bf16 value.
pip install -U "glyd[gpu]" # or: brew upgrade glyd
glyd pack and glyd verify
The glyd command packs a bf16 checkpoint on the CPU into the files python -m glyd.gpu pack writes, byte for byte, in either layout, and glyd verify checks every tensor’s sha256. On 16 CPU cores, Qwen3-8B’s 16.4 GB packed in 18.5–23.0 s in mma and 22.8–24.6 s in mma12 (the log). They are commands of glyd-gpu, the GPU weights’ own program, which now ships beside glyd; Homebrew installs it too. How to use them: getting started.
12-bit saves, loaded as saved
glyd.save_pretrained(model, path, layout="mma12"), python -m glyd.gpu pack ... --layout mma12 and glyd pack ... --layout mma12 write glyd-v3, the packs as an A10, A100 or H100 runs them, and from_pretrained loads them with no packing again: Qwen3-8B in 1.17–1.18 s against 3.54–3.61 s from its bf16 checkpoint, on an RTX 4080 SUPER with layout="mma12" (the log).
A faster 12-bit layout
The 12-bit layout is now faster: the same size and the same bits, so every product is bit for bit what it was. A layer’s time against v0.24.0’s 12-bit layout:
| GPU | Tokens | Time against v0.24.0’s, × |
|---|---|---|
| H100 SXM | 32–1,024 | 0.933–0.969 |
| H100 SXM | 1–16 | 0.986–0.992 |
| H100 PCIe | 32–1,024 | 0.942–0.991 |
| A100 | 64–128 | 0.923–0.987 |
| A100 | 256–768 | 0.969–0.982 |
| A10, L4 | 1–1,024 | 0.989–1.010 |
The logs. Exact mode and the longest prompts take 3.0–5.8% longer on an H100 SXM and 1.0–1.6% on an A10; a fix is under way (the known gaps).
Pack 12-bit models again. The 12-bit layout’s bytes changed from v0.24.0’s, and the library refuses v0.24.0’s 12-bit packs: an engine that kept them, through the C API, packs those models again. mma saves are unchanged.
The CUDA library: C API 5
- Routes. The library chooses a product’s kernel itself (
glyd_gpu_mma_route,glyd_gpu_mma12_route), so C, Rust and Python route alike, andglyd_gpu_mma_linearandglyd_gpu_mma12_linearrun it on every GPU, Hopper’s included. - The one refusal. Where a matrix’s K is not a multiple of 64, past 64 tokens,
linearreturnscudaErrorNotSupportedand launches nothing: decode the matrix and multiply by a GEMM of your own. - Older 12-bit packs: v0.24.0’s are refused with
cudaErrorInvalidValue. - Rust: the
glyd-gpucrate over the C API, in the repository.
The details: the CUDA library. Where Glyd is still slower than bf16: the known gaps. Every change, with its numbers: the changelog and the release notes.