Glyd

Release · Sep 29, 2026 · Glyd

Glyd v0.25.0: glyd pack and verify in Rust, 12-bit saves, a faster 12-bit layout

Glyd v0.25.0 packs and checks a model on the CPU with glyd pack and glyd verify, loads 12-bit saves as saved, and runs the 12-bit layout faster.

Glyd v0.25.0 packs and checks a model from the command line, with no Python or GPU; saves models in the 12-bit layout and loads them as saved; and runs the 12-bit layout faster, every product bit for bit the same. Every weight still decodes to its exact bf16 value.

pip install -U "glyd[gpu]"          # or: brew upgrade glyd

glyd pack and glyd verify

The glyd command packs a bf16 checkpoint on the CPU into the files python -m glyd.gpu pack writes, byte for byte, in either layout, and glyd verify checks every tensor’s sha256. On 16 CPU cores, Qwen3-8B’s 16.4 GB packed in 18.5–23.0 s in mma and 22.8–24.6 s in mma12 (the log). They are commands of glyd-gpu, the GPU weights’ own program, which now ships beside glyd; Homebrew installs it too. How to use them: getting started.

12-bit saves, loaded as saved

glyd.save_pretrained(model, path, layout="mma12"), python -m glyd.gpu pack ... --layout mma12 and glyd pack ... --layout mma12 write glyd-v3, the packs as an A10, A100 or H100 runs them, and from_pretrained loads them with no packing again: Qwen3-8B in 1.17–1.18 s against 3.54–3.61 s from its bf16 checkpoint, on an RTX 4080 SUPER with layout="mma12" (the log).

A faster 12-bit layout

The 12-bit layout is now faster: the same size and the same bits, so every product is bit for bit what it was. A layer’s time against v0.24.0’s 12-bit layout:

GPU Tokens Time against v0.24.0’s, ×
H100 SXM 32–1,024 0.933–0.969
H100 SXM 1–16 0.986–0.992
H100 PCIe 32–1,024 0.942–0.991
A100 64–128 0.923–0.987
A100 256–768 0.969–0.982
A10, L4 1–1,024 0.989–1.010

The logs. Exact mode and the longest prompts take 3.0–5.8% longer on an H100 SXM and 1.0–1.6% on an A10; a fix is under way (the known gaps).

Pack 12-bit models again. The 12-bit layout’s bytes changed from v0.24.0’s, and the library refuses v0.24.0’s 12-bit packs: an engine that kept them, through the C API, packs those models again. mma saves are unchanged.

The CUDA library: C API 5

  • Routes. The library chooses a product’s kernel itself (glyd_gpu_mma_route, glyd_gpu_mma12_route), so C, Rust and Python route alike, and glyd_gpu_mma_linear and glyd_gpu_mma12_linear run it on every GPU, Hopper’s included.
  • The one refusal. Where a matrix’s K is not a multiple of 64, past 64 tokens, linear returns cudaErrorNotSupported and launches nothing: decode the matrix and multiply by a GEMM of your own.
  • Older 12-bit packs: v0.24.0’s are refused with cudaErrorInvalidValue.
  • Rust: the glyd-gpu crate over the C API, in the repository.

The details: the CUDA library. Where Glyd is still slower than bf16: the known gaps. Every change, with its numbers: the changelog and the release notes.