Release · Sep 28, 2026 · Glyd
Glyd v0.23.0: the GPU kernels as a CUDA library, faster prompts on RTX 40
Glyd v0.23.0 ships its GPU kernels as a CUDA library for engines in C, C++ and Rust, with a C header and example, and runs prompts faster on RTX 40 GPUs.
Glyd v0.23.0 ships its GPU kernels as a CUDA library of their own, for engines written in C, C++, Rust or any language with a C FFI, with no Python and no PyTorch. It also multiplies prompts faster on RTX 40-series GPUs, with the same bits.
The CUDA library
Every release now carries the kernels for Linux x86_64 and aarch64 and for CUDA 12 and 13, as glyd-gpu-TAG-linux-ARCH-cudaN.tar.gz, each with its .sha256. Inside: the library, which carries its own CUDA runtime and needs only the NVIDIA driver; glyd_gpu.h, the C header that declares its 37 functions (C API version 2); an example in C; and its license.
gcc -O2 -I . -I /usr/local/cuda/include unpack.c -o unpack -L . -lglyd_gpu_cuda13 -L /usr/local/cuda/lib64 -lcudart -Wl,-rpath,"$PWD:/usr/local/cuda/lib64"
The example reads a matrix of a model saved by glyd.save_pretrained, decodes it on the GPU with the library and checks it against the bf16 checkpoint it was packed from. On an RTX 4080 SUPER, every one of Qwen3-0.6B’s 112 packs, its 196 Linears with q, k, v and gate, up merged, decodes to the checkpoint’s bits. The downloads, the header, and calling it from C++ or Rust: the CUDA library.
Faster prompts on RTX 40
Prompts on RTX 40-series GPUs are faster, and the outputs are bit for bit what they were. One forward pass on an RTX 4080 SUPER in the 12-bit layout, q, k, v and gate, up merged:
| Prompt, ms | bf16 | Glyd v0.23.0 | Glyd v0.22.0 |
|---|---|---|---|
| Qwen3-4B-Instruct-2507 | |||
| 256 tokens | 28.1 | 28.3 | 28.9 |
| 384 tokens | 39.5 | 40.8 | 41.9 |
| 512 tokens | 49.1 | 50.9 | 51.8 |
| 640 tokens | 62.7 | 64.2 | 65.8 |
| 768 tokens | 73.0 | 75.2 | 76.6 |
| Qwen3-1.7B | |||
| 384 tokens | 18.7 | 18.6 | 19.2 |
| 640 tokens | 25.9 | 26.8 | 27.8 |
| 1,024 tokens | 42.3 | 42.5 | 43.5 |
| 1,536 tokens | 63.0 | 63.6 | 67.8 |
Generating for up to 64 sequences is unchanged. A step of 65 sequences or more is faster, with the same bits:
| Tokens/s at 128 sequences | bf16 | Glyd v0.23.0 | Glyd v0.22.0 |
|---|---|---|---|
| Qwen3-4B-Instruct-2507, 12-bit layout | 4,702 | 5,069 | 4,977 |
| Qwen3-4B-Instruct-2507, smallest layout | 4,702 | 4,971 | 4,950 |
| Qwen3-1.7B, 12-bit layout | 8,452 | 8,845 | 8,699 |
| Qwen3-1.7B, smallest layout | 8,452 | 8,384 | 8,368 |
Installed the way users install it
Every release is now installed on fresh machines the ways users install it, and checked: PyPI, with and without the gpu extra, on Linux x86_64, Linux aarch64 and macOS arm64; Homebrew on macOS and Linux; and the release’s own files against their checksums (install-check.yml).
What is not there yet
Serving through vLLM is next (the plan), and the KV cache compression is in the harness only. Where Glyd is still slower than bf16: the known gaps. Every change, with its numbers: the changelog.
The license
The library, like all of Glyd’s GPU code, is source-available under the Business Source License 1.1: free for personal, educational, research and other non-commercial use, commercial production use needs a license, and each version becomes Apache-2.0 four years after its release. The codec is BSD-3-Clause OR GPL-2.0, as zstd is (the licenses).