Glyd

Release · Sep 28, 2026 · Glyd

Glyd v0.23.0: the GPU kernels as a CUDA library, faster prompts on RTX 40

Glyd v0.23.0 ships its GPU kernels as a CUDA library for engines in C, C++ and Rust, with a C header and example, and runs prompts faster on RTX 40 GPUs.

Glyd v0.23.0 ships its GPU kernels as a CUDA library of their own, for engines written in C, C++, Rust or any language with a C FFI, with no Python and no PyTorch. It also multiplies prompts faster on RTX 40-series GPUs, with the same bits.

The CUDA library

Every release now carries the kernels for Linux x86_64 and aarch64 and for CUDA 12 and 13, as glyd-gpu-TAG-linux-ARCH-cudaN.tar.gz, each with its .sha256. Inside: the library, which carries its own CUDA runtime and needs only the NVIDIA driver; glyd_gpu.h, the C header that declares its 37 functions (C API version 2); an example in C; and its license.

gcc -O2 -I . -I /usr/local/cuda/include unpack.c -o unpack -L . -lglyd_gpu_cuda13 -L /usr/local/cuda/lib64 -lcudart -Wl,-rpath,"$PWD:/usr/local/cuda/lib64"

The example reads a matrix of a model saved by glyd.save_pretrained, decodes it on the GPU with the library and checks it against the bf16 checkpoint it was packed from. On an RTX 4080 SUPER, every one of Qwen3-0.6B’s 112 packs, its 196 Linears with q, k, v and gate, up merged, decodes to the checkpoint’s bits. The downloads, the header, and calling it from C++ or Rust: the CUDA library.

Faster prompts on RTX 40

Prompts on RTX 40-series GPUs are faster, and the outputs are bit for bit what they were. One forward pass on an RTX 4080 SUPER in the 12-bit layout, q, k, v and gate, up merged:

Prompt, ms bf16 Glyd v0.23.0 Glyd v0.22.0
Qwen3-4B-Instruct-2507
256 tokens 28.1 28.3 28.9
384 tokens 39.5 40.8 41.9
512 tokens 49.1 50.9 51.8
640 tokens 62.7 64.2 65.8
768 tokens 73.0 75.2 76.6
Qwen3-1.7B
384 tokens 18.7 18.6 19.2
640 tokens 25.9 26.8 27.8
1,024 tokens 42.3 42.5 43.5
1,536 tokens 63.0 63.6 67.8

Generating for up to 64 sequences is unchanged. A step of 65 sequences or more is faster, with the same bits:

Tokens/s at 128 sequences bf16 Glyd v0.23.0 Glyd v0.22.0
Qwen3-4B-Instruct-2507, 12-bit layout 4,702 5,069 4,977
Qwen3-4B-Instruct-2507, smallest layout 4,702 4,971 4,950
Qwen3-1.7B, 12-bit layout 8,452 8,845 8,699
Qwen3-1.7B, smallest layout 8,452 8,384 8,368

Installed the way users install it

Every release is now installed on fresh machines the ways users install it, and checked: PyPI, with and without the gpu extra, on Linux x86_64, Linux aarch64 and macOS arm64; Homebrew on macOS and Linux; and the release’s own files against their checksums (install-check.yml).

What is not there yet

Serving through vLLM is next (the plan), and the KV cache compression is in the harness only. Where Glyd is still slower than bf16: the known gaps. Every change, with its numbers: the changelog.

The license

The library, like all of Glyd’s GPU code, is source-available under the Business Source License 1.1: free for personal, educational, research and other non-commercial use, commercial production use needs a license, and each version becomes Apache-2.0 four years after its release. The codec is BSD-3-Clause OR GPL-2.0, as zstd is (the licenses).