Glyd

Release · Oct 3, 2026 · Glyd

Glyd v0.27.0: lossless KV cache in vLLM

Glyd v0.27.0 holds vLLM's KV cache in fewer bits, bit for bit: 1.25 to 1.30 times the tokens on five models. The GPU package now ships compiled.

Glyd v0.27.0 holds vLLM’s KV cache in fewer bits and gives every key and value back bit for bit. On the five models measured it holds 1.25 to 1.30 times vLLM’s tokens in the same memory, and on an A100, an L4, an H100 and a GH200 decoding is faster than with vLLM’s own cache at every batch measured. The GPU parts now ship compiled, as the glyd-gpu package, with the same install commands and the same license.

The lossless KV cache

The weights were v0.26.0’s; the cache is v0.27.0’s. By default it is held on the GPUs it was measured on, and vLLM’s own cache stays everywhere else:

vllm serve Qwen/Qwen3-8B --quantization glyd                   # kv auto, the default
GLYD_KV=lossless vllm serve Qwen/Qwen3-8B --quantization glyd  # on any GPU it supports
GLYD_KV=off vllm serve Qwen/Qwen3-8B --quantization glyd       # vLLM's own cache
GLYD_KV=auto glyd run Qwen/Qwen3-8B                            # glyd run keeps vLLM's own unless GLYD_KV is set
  • auto, the default, holds the cache on an A100, an L4, an H100 and a GH200 (measured on an A100 SXM4 40 GB, an L4, an H100 SXM and a GH200), and leaves vLLM’s own on any other GPU with a line in the log that says so: an L40S, an RTX 40, an H200 and an A10 are not measured yet. GLYD_KV or --additional-config '{"glyd": {"kv": ...}}' sets it: auto, lossless or off.
  • glyd run keeps vLLM’s own cache, one user’s chat, because of the first start: the cache is set up once for the model, kept in ~/.cache/glyd/kv: on an L4, 130 s for an 8,192-token window and 217 s for Qwen3-8B’s whole 40,960-token window. glyd serve and vllm serve --quantization glyd use auto. A window past 40,960 tokens keeps vLLM’s own cache.
  • Refused, with why, for lossless: tensor, pipeline and context parallel, speculative decoding, sliding-window and linear-attention layers and the rest of what the docs list.

Measured

Qwen3-8B, bf16 weights, vLLM 0.30.0 in its default mode, against vLLM’s own cache. The KV cache in tokens, and decode tokens a second at the same batch, 128 new tokens a request:

GPUKV cache, tokensDecode, ×
vLLM’s ownLossless×LowestHighest
L4 log32,20841,4081.2861.0071.053
A100 SXM4 40 GB log136,592177,3921.2991.0181.032
H100 SXM log382,448499,3121.3061.0051.197
GH200 log482,864630,1121.3051.0031.187

On the H100, with each cache at the largest batch it holds (46 requests of 8,192 tokens with vLLM’s own, 60 with the lossless cache), decoding made 1.342 times the tokens a second; on the GH200 (58 and 76 requests), 1.305. Every batch, with its numbers: the benchmarks.

With vllm bench serve, 1,024 tokens in and 256 out, saturated, every pass with prompts of its own, the lossless cache alone served 1.23 times vLLM’s requests a second on an L4 (0.93 against 0.76), 1.29 on an A100 (7.62 against 5.92), 1.05 on an H100 SXM (21.68 against 20.58) and 1.06 on a GH200 (22.47 against 21.18). On the L4 with Glyd’s weights as well, against bf16 and vLLM’s own cache, it served 1.59 times the requests a second (1.20 against 0.76) with 2.64 times the KV tokens (71,248 against 27,024), its first token after 0.58 times the time and each token 1.55 times as long (the log).

On five models, an L4, --max-model-len 8704, the KV cache held 1.2982 times vLLM's tokens (Qwen3-4B), 1.2834 times vLLM's tokens (Qwen3-8B), 1.2773 times vLLM's tokens (Qwen2.5-7B), 1.2743 times vLLM's tokens (Mistral-7B-v0.3), 1.2526 times vLLM's tokens (Llama-3.1-8B) (the log).

Where it costs

With long prompts on an H100 SXM, 8,192 tokens in and 256 out, saturated, it served 1.05 times vLLM’s requests a second (2.80 against 2.67) and its first token came after 0.89 times the time (8,196 ms against 9,199), but each generated token took 1.12 times as long (49.2 ms against 43.9). The KV cache held 494,256 tokens against 377,024 (1.31 times) and the model’s memory was 1.07 times (16.34 against 15.27 GiB) (the log). On a GH200 the same run served 1.12 times the requests a second (3.12 against 2.78), with its first token after 0.95 times the time and each token 0.97 times as long (the log). An H200 is not measured.

Prompts of a model’s rare tokens, tokens whose embeddings were never trained, are the cache’s known limit: 40 of them of 1,400 tokens at once with 8 ordinary prompts were all served on every one of five models, at 0.88 times (Llama-3.1-8B) and 0.77 times (Mistral-7B-v0.3) vLLM’s prompt tokens a second, and at 1.00 to 1.01 times on the Qwen models, which have none (the known gaps).

What is exact

The stored values. Every value read back was the value written: 0 of 25,683,296,256 differ in a stress run. Prompt logprobs are vLLM’s bit for bit, a 36,000-token prompt’s 35,999 of 35,999. A decode step’s attention is not bit-equal to vLLM’s, so greedy tokens can part from vLLM’s after some tokens, as they do with vLLM’s own attention when one of its settings changes; with exact, which runs vLLM’s own attention for decode steps too, the 14 prompts of a prefix-hit scenario on an A100 were bit for bit. The model checks against vLLM’s own pass on all five models (the docs).

The GPU parts ship compiled

The GPU half of Glyd is now the glyd-gpu package, compiled wheels for Linux x86_64 and aarch64, which pip install "glyd[gpu]" and pip install "glyd[vllm]" bring; curl -LsSf https://getglyd.com/install.sh | sh and glyd run MODEL are the same. The license is unchanged, the Business Source License 1.1 on the same terms (the licenses); versions up to v0.26.0 stay source-available in their tags. The C API’s public header, glyd_gpu.h, comes in the wheel and carries 8 functions, version 14 (the CUDA library). glyd pack and glyd verify are the Python tool’s commands, as glyd run is; the release’s files and Homebrew carry glyd and glyd-store.

Fixes: with transformers 5.18, Inkling’s embedding is packed with its norm kept and several embedding classes that are not plain lookups are no longer packed; glyd run with the KV cache no longer cuts the window to fit it.

Where Glyd is still slower than bf16: the known gaps. Every change, with its numbers: the changelog and the release notes.