Release · Sep 28, 2026 · Glyd
Glyd v0.22.0: mixtures of experts saved packed, faster on H100 and A100
Glyd v0.22.0 saves and loads mixtures of experts packed, packs Llama 4's experts, and takes less GPU time than bf16 on an H100 up to 64 sequences a step.
Glyd v0.22.0 saves a mixture of experts with its experts packed and loads it back, packs the experts of every mixture-of-experts family in transformers 5.17, Llama 4’s among them, and takes less time on H100 and A100 GPUs. Every weight still decodes to its exact bf16 value.
pip install -U "glyd[gpu]"
Mixtures of experts, saved and loaded
glyd.save_pretrained now saves a mixture of experts with its experts packed, and glyd.from_pretrained loads it without packing it again; verify=True checks every tensor, and python -m glyd.gpu pack and verify take one. Such a checkpoint is in the glyd-v2 format, which glyd 0.21 refuses; a dense model’s stays glyd-v1. granite-3.1-3b-a800m on an RTX 4080 SUPER, its logits bit for bit those of the model packed as it loads:
| granite-3.1-3b-a800m | bf16 checkpoint | Saved packed |
|---|---|---|
| Safetensors, GB | 6.60 | 4.66 |
| Loading, s | 3.9 | 0.4 |
Loading with verify=True, s |
3.4 |
Every mixture-of-experts family of transformers 5.17 now has its experts packed, those whose own code runs them too: Llama 4, DBRX, Aria, JetMoE, Step 3.7 and LongCat-Flash. Llama 4’s experts had stayed bf16. A Llama-4-Scout-17B-16E-Instruct MoE block at its real sizes on an RTX 4080 SUPER, bit for bit with exact=True:
| Llama-4-Scout-17B-16E-Instruct MoE block | bf16 | Glyd |
|---|---|---|
| 1 token, ms | 6.2 | 0.7 |
| 512 tokens, ms | 23.4 | 6.4 |
| Experts, GB | 4.03 | 2.70 |
Faster on H100 and A100
On Hopper, steps of 17 to 128 tokens and prompts of 129 to 512 tokens now take less time. A step’s GPU time on an H100 PCIe, q, k, v and gate, up merged for bf16 and Glyd alike:
| GPU time a step, ms | Sequences a step | |||
|---|---|---|---|---|
| 1 | 8 | 32 | 64 | |
| Qwen3-8B, bf16 | 12.30 | 13.34 | 14.57 | 15.66 |
| Qwen3-8B, Glyd | 11.13 | 12.32 | 13.48 | 14.64 |
| Qwen3-32B, bf16 | 42.44 | 44.40 | 47.10 | 49.56 |
| Qwen3-32B, Glyd | 34.65 | 37.34 | 41.24 | 44.41 |
Prompts on the H100 PCIe, and Qwen3-8B’s step on an A100 in the 12-bit layout:
| bf16 | Glyd v0.22.0 | Glyd v0.21.0 | |
|---|---|---|---|
| H100 PCIe, Qwen3-32B’s prompt, ms | |||
| 256 tokens | 70.3 | 79.7 | 142.1 |
| 512 tokens | 113.1 | 141.0 | 189.8 |
| A100, Qwen3-8B’s step, ms | |||
| 32 sequences | 21.12 | 18.54 | 19.14 |
| 64 sequences | 21.13 | 21.05 | 21.69 |
| 128 sequences | 25.71 | 28.17 | 29.45 |
Prompts on an RTX 4080 SUPER in the smallest layout (mma):
| Prompt, ms | bf16 | Glyd |
|---|---|---|
| Qwen3-1.7B, 128 tokens | 10.8 | 10.3 |
| Qwen3-1.7B, 256 tokens | 14.0 | 14.6 |
| Qwen3-1.7B, 512 tokens | 24.3 | 23.8 |
| Qwen3-4B-Instruct-2507, 2,048 tokens | 198.3 | 199.5 |
| Qwen3-4B-Instruct-2507, 4,096 tokens | 445.4 | 447.5 |
The prebuilt library now carries native code for Blackwell, sm_100 and sm_120.
What is still slower
v0.22.0 ran on six GPUs, bf16 and Glyd on each (tested GPUs). At 128 sequences a step, and for long prompts on most of them, Glyd still took longer than bf16: the known gaps. Every change, with its numbers: the changelog.