Release · Sep 28, 2026 · Glyd
Glyd v0.24.0: generate() compiled by default, faster Hopper prompts
Glyd v0.24.0 compiles generate() by default on PyTorch 2.13 or later, and runs Hopper's prompts and an A10's long prompts faster.
Glyd v0.24.0 runs generate() compiled by default, with no code change: on an RTX 4080 SUPER, Qwen3-1.7B generates about twice the tokens a second. It also runs Hopper’s prompts faster, fixes the CUDA 12 library on Hopper, and runs an A10’s long prompts faster. Every weight still decodes to its exact bf16 value.
pip install -U "glyd[gpu]"
generate(), compiled by default
On PyTorch 2.13 or later, generate() on a model from glyd.from_pretrained runs compiled, as generate(..., cache_implementation="static") asks transformers to run it: a static cache and the forward pass in a CUDA graph. Plain generate(), 128 new tokens, on an RTX 4080 SUPER:
| Tokens/s | 1 sequence | 8 sequences | ||
|---|---|---|---|---|
| v0.23.0 | v0.24.0 | v0.23.0 | v0.24.0 | |
| Qwen3-1.7B | 96.6 | 184.6 | 772.1 | 1,260.4 |
| Qwen3-4B-Instruct-2507 | 75.1 | 94.3 | 562.8 | 609.6 |
| Qwen3-8B | 48.6 | 55.3 | 365.2 | 385.8 |
| granite-3.1-3b-a800m-instruct | 90.3 | 232.3 | 698.9 | 1,546.5 |
- The first call compiles. Qwen3-8B’s took 17.5 s with PyTorch’s compile caches empty, 6.7 s in a later process: warm up with one short
generate()before serving. exact=Trueis never compiled and stays bit-identical to bf16.- To run it eager, as before:
compile=Falseinglyd.from_pretrainedorglyd.gpu.compress, orGLYD_COMPILE=0; for one call,generate(..., disable_compile=True). Below PyTorch 2.13 it stays eager. - Greedy tokens compiled can differ from the eager loop’s, as a compiled bf16 model’s can from its eager ones.
What else runs eager, and when: getting started.
Prompts on Hopper
Prompts of 129 to 1,024 tokens on Hopper are faster. On an H100 SXM, Qwen3-8B’s forward pass over 1,024 tokens takes 45.0 ms against bf16’s 36.9, still slower (the known gaps); work on datacenter prompts is under way. The log
The CUDA 12 library, which the wheels load for PyTorch built for CUDA 12, had a slower path on Hopper, in v0.22.0 and v0.23.0 too. It no longer does: its products now take as long as the CUDA 13 library’s, with the same outputs, bit for bit.
Long prompts on an A10
An A10’s long prompts are faster. Qwen3-8B’s forward pass, over bf16’s time:
| Prompt, tokens | v0.23.0, % over bf16 | v0.24.0, % over bf16 |
|---|---|---|
| 1,024 | +30.3 | +10.0 |
| 2,048 | +38.8 | +5.2 |
| 4,096 | +50.9 | +2.6 |
The C API
The CUDA library’s C API is version 3 (the CUDA library).
Where Glyd is still slower than bf16: the known gaps. Every change, with its numbers: the changelog and the release notes.