Release · Sep 29, 2026 · Glyd
Glyd v0.25.1: long prompts on the L4 and L40S, faster on Hopper
Glyd v0.25.1 brings long prompts on the L4 and L40S within a few percent of bf16's at 8,192 tokens, and makes exact mode on Hopper faster again.
Glyd v0.25.1 runs long prompts on NVIDIA’s L4 and L40S close to bf16’s time, and makes exact mode and the longest prompts on Hopper faster again. The same bits: every weight still decodes to its exact bf16 value.
pip install -U "glyd[gpu]" # or: brew upgrade glyd
Long prompts on the L4 and L40S
At their power caps, these cards were slower on long prompts. Long prompts on them now run faster; short prompts are unchanged. Qwen3-8B’s forward pass, smallest layout, over bf16’s time in the same run:
| Prompt, tokens | L4, % over bf16 | L40S, % over bf16 | ||
|---|---|---|---|---|
| v0.25.0 | v0.25.1 | v0.25.0 | v0.25.1 | |
| 1,024 | +27.8 | +25.5 | +38.1 | +30.3 |
| 2,048 | +31.0 | +11.3 | +41.2 | +11.9 |
| 4,096 | +37.9 | +8.4 | +39.4 | +10.9 |
| 8,192 | +98.4 | +4.7 | +35.0 | +3.8 |
mma, the smallest layout, stays these cards’ default, for 33% less memory. For the fastest short prompts, at 25% less, load with layout="mma12": its prompts took 5 to 20% less time than mma’s to 1,536 tokens on an L4, and 4.0 to 18.1% less from 512 to 1,536 tokens on an L40S.
Faster on Hopper
On Hopper, exact mode and the longest prompts in the 12-bit layout are faster again. On a GH200 they take 2.5 to 3.6% less time than before v0.25.0, where v0.25.0’s took 3.5 to 6.0% more; the H100 SXM was not run again (the log). Every other GPU keeps v0.25.0’s behavior.
Where Glyd is still slower than bf16: the known gaps. Every change, with its numbers: the changelog and the release notes.