The CUDA library
From v0.27.0 Glyd’s GPU kernels are a library of their own in the glyd-gpu wheel, with the C header that declares it: a bf16 model’s weights held compressed in GPU memory, with functions that unpack them on the GPU bit for bit and multiply by them. It is for engines in C, C++, Rust or any language with a C FFI, with no Python and no PyTorch. The Python package calls the same library (getting started); this page is for calling it from your own code. Releases v0.23.0 to v0.26.0 carried the library as downloads of their own, with a larger header: below.
In the glyd-gpu wheel
pip install glyd-gpu # or pip install "glyd[gpu]", which brings it
python -c "import glyd_gpu, os; print(os.path.dirname(glyd_gpu.__file__))"
The directory it prints holds libglyd_gpu_cuda12.so and libglyd_gpu_cuda13.so, the library for CUDA 12 and 13, and glyd_gpu.h, the C header; the wheel’s LICENSE is the library’s license (below). The libraries carry their own CUDA runtime, so they need only the NVIDIA driver and the system’s C and C++ libraries, and hold code for Ampere, Ada, Hopper and Blackwell. A GPU the library has no code for gets cudaErrorNotSupported, with nothing launched. Linking:
gcc prog.c -I DIR -I /usr/local/cuda/include -L DIR -lglyd_gpu_cuda13 -L /usr/local/cuda/lib64 -lcudart -Wl,-rpath,DIR
The header, glyd_gpu.h, carries 8 functions and GLYD_GPU_API_VERSION (14), for a matrix W [O, K] of a model saved by glyd pack (or glyd.save_pretrained), passed as the save holds it: its arrays, copied to device memory, and its parameters, as glyd.json gives them, in host memory. glyd.json’s layout, mma or mma12, says which functions take it.
- The version.
glyd_gpu_api_version()againstGLYD_GPU_API_VERSION: the library’s functions are the header’s where the two agree, and a C FFI does not see a call’s arguments, so check it.glyd_gpu_error_string(status)gives a status’s text. - A product.
glyd_gpu_mma_linearandglyd_gpu_mma12_linear: Y [M, O] = X Wᵀ (+ bias [O], or NULL), for X [M, K] and Y [M, O] row-major bf16;modeisGLYD_GPU_AUTO(-1), the library running it the way it measured fastest on the current GPU for M tokens. Where the library has no kernel for a W and M, the call returnscudaErrorNotSupported: unpack W and multiply by a GEMM of your own. - Its workspace. First
glyd_gpu_mma_linear_workspaceorglyd_gpu_mma12_linear_workspace(O, K, M, mode, &bytes), on the device the product will run on, then the product with a buffer of at least those bytes (NULL where 0) andGLYD_GPU_DONE_COUNTERS(O, M)int32 counters in device memory, zero before its first call and left zero by each. A workspace and a set of counters serve one stream at a time. - An unpack.
glyd_gpu_mma_unpackandglyd_gpu_mma12_unpackgive W back to bf16 bit for bit, intoout[O, K]. - Every call. The arrays are in device memory, 16-byte aligned (
cudaMalloc’s are; bf16 as its bits,uint16_t), but for the parameters and a workspace query’s bytes, in host memory; sizes are values; the kernels are launched on the stream you pass (0 for the default stream) of the current device, and run after the call returns, as any launch. Every output, workspace and counter is yours. - The return. 0, or a
cudaError_t:cudaErrorInvalidValuefor an argument out of range,cudaErrorNotSupportedwhere the library has no kernel for the call on the current GPU (nothing launched), else the launch’s.
The downloads, v0.23.0 to v0.26.0
The last release that carried the library as downloads is v0.26.0. Take the download for your machine and for the CUDA major version of your toolkit and runtime: cuda12 for CUDA 12 (built with 12.8), cuda13 for CUDA 13.
| Linux | CUDA | Download | Checksum |
|---|---|---|---|
| x86_64 | 12 | glyd-gpu-v0.26.0-linux-x86_64-cuda12.tar.gz | .sha256 |
| x86_64 | 13 | glyd-gpu-v0.26.0-linux-x86_64-cuda13.tar.gz | .sha256 |
| aarch64 | 12 | glyd-gpu-v0.26.0-linux-aarch64-cuda12.tar.gz | .sha256 |
| aarch64 | 13 | glyd-gpu-v0.26.0-linux-aarch64-cuda13.tar.gz | .sha256 |
Each unpacks to a folder of five files (the header described below is v0.26.0’s):
libglyd_gpu_cuda12.soorlibglyd_gpu_cuda13.so: the kernels. The library carries its own CUDA runtime, linked in, so it needs only the NVIDIA driver and the system’s C and C++ libraries (glibc 2.28 or later). It holds code for Ampere (sm_80, sm_86), Ada (sm_89), Hopper (sm_90a) and Blackwell (sm_100, sm_120), and PTX for the GPUs after them.glyd_gpu.h: the C API.unpack.c: an example in C alone.LICENSE: the library’s license, gpu/LICENSE (below).README.md: the same, with the example’s build line for that download’s library.
Check a download against its checksum before you use it:
sha256sum -c glyd-gpu-*-linux-x86_64-cuda13.tar.gz.sha256
tar -xzf glyd-gpu-*-linux-x86_64-cuda13.tar.gz
Every release is also installed and checked the way a user installs it, on a fresh machine (install-check.yml).
v0.26.0’s header
glyd_gpu.h of v0.26.0 declares the C API: its 52 functions, GLYD_GPU_API_VERSION (7), the arrays of each packed layout, the workspace queries, the stream and the return codes. The library’s own build includes it, so each function is held to its declaration, and the Python package’s calls are checked against it.
- The version.
glyd_gpu_api_version()againstGLYD_GPU_API_VERSION: the library’s functions are the header’s where the two agree, and a C FFI does not see a call’s arguments, so check it. - A call. The arrays are in device memory (bf16 as its bits,
uint16_t) but for a layout’s few words and a workspace query’s bytes, in host memory; the sizes are values; and the kernels are launched on the stream you pass (0 for the default stream) of the current device, running after the call returns, as any launch. Every output, workspace and counter is yours. - A product with a workspace. First
glyd_gpu_NAME_workspace(its sizes, &bytes), on the device it will run on, thenglyd_gpu_NAMEwith a buffer of at least those bytes. A workspace and a set of counters serve one stream at a time. - A product, as Glyd routes it.
glyd_gpu_mma_routeandglyd_gpu_mma12_routechoose the route of a product of M tokens on a GPU (glyd_gpu_gpu: the current one’s code), as the Python package does;glyd_gpu_mma_linearandglyd_gpu_mma12_linearrun it, Y = X Wᵀ (+ bias), on every GPU, Hopper’s included. Environment variables listed inglyd_gpu.hmove the routes’ thresholds. - The one refusal. Where K is not a multiple of 64, past 64 tokens (in the 12-bit layout also from
GLYD_DEC_MINtokens where that is lower),linearand its workspace query returncudaErrorNotSupportedand launch nothing: decode the matrix (glyd_gpu_*_unpack) and multiply by a GEMM of your own, as the Python package does. - Older 12-bit packs. v0.24.0’s 12-bit packs are refused with
cudaErrorInvalidValue: pack those models again. - The return. 0, or a
cudaError_t:cudaErrorInvalidValuefor an argument out of range,cudaErrorNotSupportedwhere a kernel is not for this GPU, else the launch’s.glyd_gpu_error_string()gives its text.
Each release’s GPU run checks every call of the library against the same kernels built from source, bit for bit, on every GPU it runs on (tested GPUs).
Long prompts on an A100 SXM, a GH200 and an H100 SXM
From v0.26.0 a long prompt on these GPUs can take a faster route, as measured: a forward pass takes 0.85 to 0.97 of v0.25.1’s time on the A100 SXM, 0.91 to 0.95 with Qwen3-32B on the GH200 and 0.89 to 0.94 with Qwen3-14B on the H100 SXM. No prompt past 8,192 tokens takes it, and an H200, an H100 NVL, the PCIe cards and a MIG slice keep v0.25.1’s routes. The times, by prompt length: v0.26.0’s release post.
- In v0.26.0’s C API, opt-in:
glyd_gpu_mma12_routegives the route only for a GPU code withGLYD_GPU_WITH_SPLITadded, which the Python package’s Linears ask for;linear’s own route (-1) never takes it, so the glyd-gpu crate and the vLLM plugin keep v0.25.1’s routes. - From C. v0.26.0’s callers find the long-prompt calls and their arguments in its
glyd_gpu.h(glyd_gpu_ring_*andglyd_gpu_mma12_ring_*); v0.27.0’s public header has no calls for the route. - Memory. In the Python package the route takes 600 MiB for Qwen3-8B, 1.0 GiB for 14B and 1.5 GiB for 32B beyond the weights, and a 32 MiB workspace.
- Its products. The same bits run to run within a process, but not bit for bit bf16’s, so
exact=Truenever takes the route. - Where it cannot run (a driver before CUDA 12.5, or a stream being captured into a CUDA graph), the calls return
cudaErrorNotSupported: take the route the code without the flag gives.GLYD_SPLIT_MIN=-1turns the route off;GLYD_SPLIT_MINandGLYD_SPLIT_MAXmove it. - The GPU codes. A code now tells an A100 PCIe (5080), an H100 PCIe (5090), a GH200 (6090) and an H100 SXM (7090) apart, by
GLYD_GPU_PCIE,GLYD_GPU_GH200andGLYD_GPU_H100; an H100 NVL and an H200 stay 90.
The example, v0.26.0’s
unpack.c of the v0.26.0 download reads a matrix of a model saved by glyd.save_pretrained (the glyd-v1 format: the packs in safetensors, their shapes in glyd.json), decodes it on the GPU with the library, and checks it against the bf16 checkpoint it was packed from, bit for bit; a pack of merged Linears (q, k, v; gate, up) tensor by tensor. Built in the download’s folder with the CUDA toolkit’s headers and runtime (/usr/local/cuda, or yours), as any C program builds against the library:
gcc -O2 -I . -I /usr/local/cuda/include unpack.c -o unpack -L . -lglyd_gpu_cuda13 -L /usr/local/cuda/lib64 -lcudart -Wl,-rpath,"$PWD:/usr/local/cuda/lib64"
With the CUDA 12 download, -lglyd_gpu_cuda12. Your program links its own CUDA runtime, for its memory and streams, and it runs beside the one inside the library on the same driver. A saved model comes from the Python package, and the bf16 checkpoint it was packed from is in the Hugging Face cache:
python -m glyd_gpu pack Qwen/Qwen3-0.6B qwen3-0.6b-glyd
./unpack qwen3-0.6b-glyd ~/.cache/huggingface/hub/models--Qwen--Qwen3-0.6B/snapshots/*/ [PACK]
On an RTX 4080 SUPER (CUDA 13.0, v0.23.0’s library, C API 2), the first pack, then a merged one:
$ ./unpack qwen3-0.6b-glyd ~/.cache/huggingface/hub/models--Qwen--Qwen3-0.6B/snapshots/*/
libglyd_gpu: C API 2, CUDA runtime 13000
model.layers.0.self_attn.o_proj: [1024, 2048], 10.86 bits a weight packed, decoded on the GPU
model.layers.0.self_attn.o_proj.weight [1024, 2048]: the checkpoint's, bit for bit
$ ./unpack qwen3-0.6b-glyd ~/.cache/huggingface/hub/models--Qwen--Qwen3-0.6B/snapshots/*/ model.layers.0.self_attn.q_proj
libglyd_gpu: C API 2, CUDA runtime 13000
model.layers.0.self_attn.q_proj: [4096, 1024], 10.79 bits a weight packed, decoded on the GPU
model.layers.0.self_attn.q_proj.weight [2048, 1024]: the checkpoint's, bit for bit
model.layers.0.self_attn.k_proj.weight [1024, 1024]: the checkpoint's, bit for bit
model.layers.0.self_attn.v_proj.weight [1024, 1024]: the checkpoint's, bit for bit
Every one of Qwen3-0.6B’s 112 packs, its 196 Linears with q, k, v and gate, up merged, decodes to the checkpoint’s bits this way, and a bit flipped in the checkpoint is found (gpu/README.md at v0.26.0).
From C++, Rust and other languages
C++ includes the header as it is; its declarations are extern "C". Other languages call the same functions through their C FFI: device pointers as raw pointers, a stream as the CUDA runtime’s or driver’s handle. v0.23.0 to v0.26.0 also had a Rust crate, glyd-gpu, over the library (in those tags); from v0.27.0 the header is the interface.
The license
The library, like the rest of the GPU package, is under the Business Source License 1.1: the wheel carries its LICENSE, and v0.26.0’s text is gpu/LICENSE. From v0.27.0 the package ships compiled; the license is unchanged. Its grant, as the license states it:
You may make production use of the Licensed Work for personal, educational, research, and other non-commercial purposes. Any commercial or revenue-generating production use, including use within a product or service offered to third parties, requires a separate commercial license from the Licensor (contact: suryakoritala@getglyd.com).
Each version becomes Apache-2.0 four years after it is first published, or on 2030-09-18, whichever comes first for that version. The codec underneath is BSD-3-Clause OR GPL-2.0 (the licenses).
Built from source
Up to v0.26.0 the library’s source is in the repository: bash gpu/build_lib.sh in a clone of that tag builds it for your nvcc’s CUDA major version (Blackwell’s code with CUDA 12.8 or later), and gpu/README.md has the example’s build line against that build. From v0.27.0 the library comes only in the glyd-gpu wheel.