Glyd

Release · Sep 28, 2026 · Glyd

Glyd v0.22.0: mixtures of experts saved packed, faster on H100 and A100

Glyd v0.22.0 saves and loads mixtures of experts packed, packs Llama 4's experts, and takes less GPU time than bf16 on an H100 up to 64 sequences a step.

Glyd v0.22.0 saves a mixture of experts with its experts packed and loads it back, packs the experts of every mixture-of-experts family in transformers 5.17, Llama 4’s among them, and takes less time on H100 and A100 GPUs. Every weight still decodes to its exact bf16 value.

pip install -U "glyd[gpu]"

Mixtures of experts, saved and loaded

glyd.save_pretrained now saves a mixture of experts with its experts packed, and glyd.from_pretrained loads it without packing it again; verify=True checks every tensor, and python -m glyd.gpu pack and verify take one. Such a checkpoint is in the glyd-v2 format, which glyd 0.21 refuses; a dense model’s stays glyd-v1. granite-3.1-3b-a800m on an RTX 4080 SUPER, its logits bit for bit those of the model packed as it loads:

granite-3.1-3b-a800m bf16 checkpoint Saved packed
Safetensors, GB 6.60 4.66
Loading, s 3.9 0.4
Loading with verify=True, s 3.4

Every mixture-of-experts family of transformers 5.17 now has its experts packed, those whose own code runs them too: Llama 4, DBRX, Aria, JetMoE, Step 3.7 and LongCat-Flash. Llama 4’s experts had stayed bf16. A Llama-4-Scout-17B-16E-Instruct MoE block at its real sizes on an RTX 4080 SUPER, bit for bit with exact=True:

Llama-4-Scout-17B-16E-Instruct MoE block bf16 Glyd
1 token, ms 6.2 0.7
512 tokens, ms 23.4 6.4
Experts, GB 4.03 2.70

Faster on H100 and A100

On Hopper, steps of 17 to 128 tokens and prompts of 129 to 512 tokens now take less time. A step’s GPU time on an H100 PCIe, q, k, v and gate, up merged for bf16 and Glyd alike:

GPU time a step, msSequences a step
183264
Qwen3-8B, bf1612.3013.3414.5715.66
Qwen3-8B, Glyd11.1312.3213.4814.64
Qwen3-32B, bf1642.4444.4047.1049.56
Qwen3-32B, Glyd34.6537.3441.2444.41

Prompts on the H100 PCIe, and Qwen3-8B’s step on an A100 in the 12-bit layout:

bf16 Glyd v0.22.0 Glyd v0.21.0
H100 PCIe, Qwen3-32B’s prompt, ms
256 tokens 70.3 79.7 142.1
512 tokens 113.1 141.0 189.8
A100, Qwen3-8B’s step, ms
32 sequences 21.12 18.54 19.14
64 sequences 21.13 21.05 21.69
128 sequences 25.71 28.17 29.45

Prompts on an RTX 4080 SUPER in the smallest layout (mma):

Prompt, ms bf16 Glyd
Qwen3-1.7B, 128 tokens 10.8 10.3
Qwen3-1.7B, 256 tokens 14.0 14.6
Qwen3-1.7B, 512 tokens 24.3 23.8
Qwen3-4B-Instruct-2507, 2,048 tokens 198.3 199.5
Qwen3-4B-Instruct-2507, 4,096 tokens 445.4 447.5

The prebuilt library now carries native code for Blackwell, sm_100 and sm_120.

What is still slower

v0.22.0 ran on six GPUs, bf16 and Glyd on each (tested GPUs). At 128 sequences a step, and for long prompts on most of them, Glyd still took longer than bf16: the known gaps. Every change, with its numbers: the changelog.