Serving with vLLM
From v0.26.0, vLLM serves a model with its weights packed by Glyd. Its Linear layers, and a mixture of experts’ experts, stay packed on the GPU, every weight its bf16 value. vLLM sizes its KV cache after the weights load, so the memory the packs save becomes KV cache: more requests at once on the same GPU. The way in is one command, glyd run; vllm serve --quantization glyd is for a vLLM setup of your own (advanced). From v0.27.0 vLLM’s KV cache can be held in fewer bits too, bit for bit (the lossless KV cache).
Local chat, like Ollama
Two commands, on Linux with an NVIDIA GPU:
curl -LsSf https://getglyd.com/install.sh | sh
glyd run Qwen/Qwen3-8B
The script installs uv if it is missing, then Glyd with vLLM 0.30 and PyTorch as one isolated tool, on a Python 3.12 that uv fetches for it: about 8 GB of disk, and no sudo, no virtual environment and no pip. It ends with glyd doctor. glyd run downloads the model (Qwen3-8B is 16.4 GB the first time), checks that this machine can run it, starts vLLM with settings it works out from the GPU, and opens a chat: type in the terminal (/bye leaves, /clear starts over, /think turns the model’s thinking on and off), or open http://localhost:8000.
What it needs:
- Linux on x86_64 (the script also takes aarch64, which was not run).
- An NVIDIA GPU of the Ampere generation or newer (RTX 30 and 40 series, A10, A100, L4, H100 and later) and its driver, 580 or newer: the CUDA 13 build of PyTorch that vLLM 0.30 installs.
- curl, and disk: the script asks for 10 GB free in your home directory and stops with a plain message under that; the model needs its own.
- No CUDA toolkit and no sudo. vLLM’s Triton builds small launchers with a C compiler when the server starts; where the machine has no gcc or clang, the script adds ziglang, a compiler from PyPI, and
glyd runhands it to vLLM.
On a Mac, or a Linux machine with no NVIDIA GPU, glyd run has nothing to run on. The script says so (glyd run needs Linux with an NVIDIA GPU) and installs what works there: the compression program (glyd FILE -o OUT), the release’s tarball checked against its sha256, in ~/.local/share/glyd/cli and linked from ~/.local/bin. Homebrew’s brew install surya-koritala/glyd/glyd is the other way. glyd run there says the same.
What the script does to your machine, all of it under your home directory (it is one function, called on its last line, so a download that is cut short runs nothing; read it first):
- uv, where there is none: uv’s own installer at a pinned version, checked against its sha256 and told to edit no shell startup file, into
~/.local/bin. - Glyd, at the release the script names, with its roughly 200 packages at the versions the acceptance run installed (listed at the end of the script): in uv’s tool directory (
~/.local/share/uv/tools/glyd) and uv’s cache (~/.cache/uv), with a link~/.local/bin/glyd.GLYD_CONSTRAINTS=nonein front ofshresolves the packages fresh, where one has been withdrawn from PyPI. - Your PATH, where
~/.local/binis not on it:uv tool update-shelladds a line to your shell’s startup file. The script says so before it does, and uv names the file. - Nothing else, and nothing of yours replaced. Where
~/.local/bin/glydis a program uv did not put there (the compression program, a pip install of glyd, one of your own), the script stops before it installs Glyd and says how to keep both. Aglydthat comes first on your PATH (Homebrew’s, cargo’s) is the one typingglydreaches, and one older than v0.26.0 takesrunfor a file name: the script says so, with the line that puts~/.local/binfirst.
glyd run, in order, stopping with a plain message and what to do where it must:
- Checks the GPU and its driver, the compiler, vLLM’s version, the model (on the Hugging Face Hub, bf16, and not gated without your token), whether the model with Glyd fits the memory free now (if not: how much it needs, what holds the GPU’s memory, and the largest model of its family that fits), the disk the download needs, and the port.
- Downloads the model, with one progress bar. A gated model (Llama, Gemma): accept its licence on its Hugging Face page, then run
glyd login. - Chooses the settings below and prints them on one line.
- Starts vLLM with its output in a log file, and shows what it is doing.
- Chats, in the terminal and at the printed address.
glyd run MODEL --prompt "Say hello" prints one answer and exits, and --context N sets the context. Any vLLM flag after a lone -- is passed on and wins over what was chosen: glyd run MODEL -- --max-model-len 4096.
The settings
- Memory: vLLM’s share of the GPU is the memory free now, less 0.55 GiB for the server’s CUDA context and 0.4 GiB left for a desktop, at most 0.92, counted from the total CUDA reports (an RTX 4080 SUPER’s is 15.57 GiB, nvidia-smi’s 15.99).
glyd runtakes only what two windows of KV cache need,glyd servethe whole share. - Context: the model’s own length, or the most the KV cache holds at that share, in multiples of 1,024. Under 4,096 the model does not fit.
- Eager mode, always (
-- --no-enforce-eagercompiles). On an L4 with Qwen3-8B, a compiled server was 2 to 3% faster in all, for 1, 4 and 8 users (21.7, 84.8 and 165.8 tokens a second against eager’s 21.2, 82.3 and 161.2), took 2 min 45 s to come up against 47 s, and needs 2.15 GiB beyond the weights where eager needs 0.5. - The layout: the plugin’s own choice for the GPU (
mma, 10.80 bits a weight, on Ada and wherever only it fits; elsemma12),mmaalso wheremma12leaves less room than an 8,192-token chat. - The sampler: PyTorch’s. FlashInfer’s compiles with nvcc at the first request that samples, and the machine has no CUDA toolkit; the tokens a second are the same (21.1 with FlashInfer’s, 21.2 with PyTorch’s, for one user).
- Tool calls and thinking, by family: Qwen3
hermesandqwen3, Qwen3 Instruct-2507 and Qwen2.5hermes, Qwen3-Coderqwen3_coder, DeepSeek-R1 distillsdeepseek_r1, Llama 3.xllama3_json, Mistralmistral. Another family chats without tool calls. - The KV cache: vLLM’s own in
glyd run, which keeps it unlessGLYD_KVis set (GLYD_KV=auto glyd run MODELholds it in fewer bits on an A100, an L4, an H100 and a GH200, after a first start of minutes);glyd serveusesauto(the lossless KV cache). - The address: 127.0.0.1, with telemetry off (below).
The line printed before loading is these, for example Settings: 10,240-token context (the most that fits), eager mode, 61% of GPU memory (14.4 GB), tool calls (hermes), thinking shown apart (qwen3); PyTorch sampler (no CUDA toolkit). The model’s weights with Glyd are counted from its config and file sizes before anything downloads (Qwen3-8B: 12.2 GB, bf16’s 16.4).
Who can reach the server
glyd run and glyd serve listen on 127.0.0.1, this computer’s own address, and answer this computer’s programs and the page they serve. A web page you open can still send requests to 127.0.0.1, so the server checks each request: a Host that is not localhost, 127.0.0.1 or [::1] is refused with 421, and an Origin that is not the server’s own with 403. curl, the OpenAI libraries and Open WebUI’s server are answered. A web app of your own on another origin is refused too; glyd serve MODEL --host 127.0.0.1 -- --allowed-origins '["http://localhost:5173"]' lets it in and turns these checks off.
glyd serve --host ADDRESS is the way onto the network, and the checks are off for an address you chose: anyone who can reach this computer can send the model prompts. Give it a key, in the environment:
VLLM_API_KEY=YOUR_KEY glyd serve Qwen/Qwen3-8B --host 0.0.0.0
The key guards /v1 only, and the traffic is plain HTTP: use a VPN or an SSH tunnel.
glyd serve, glyd doctor and Open WebUI
glyd serve MODEL does the same checks, chooses the same settings, and leaves the server up for other programs: the OpenAI API at http://localhost:8000/v1 and the chat page at http://localhost:8000. glyd doctor prints the GPU, driver, CUDA, free memory, compilers and versions, and which of Qwen3-8B, 14B and 32B fit, at Glyd’s sizes.
Open WebUI is a chat page with accounts, history and tools that talks to that API and uses no GPU memory. It is large (its Docker image is 6.5 GB, and uvx puts 7.1 GB into uv’s cache), and pinned to 0.11.4, the version tested. Without Docker:
DATA_DIR="$HOME/.open-webui" OPENAI_API_BASE_URL=http://127.0.0.1:8000/v1 OPENAI_API_KEY=none ENABLE_PERSISTENT_CONFIG=False \
CORS_ALLOW_ORIGIN='http://localhost:3000;http://127.0.0.1:3000' \
uvx --python 3.11 open-webui@0.11.4 serve --host 127.0.0.1 --port 3000
With Docker Engine on Linux, whose containers can share the host’s network:
docker run -d --name open-webui --network=host -e PORT=3000 -e HOST=127.0.0.1 \
-e OPENAI_API_BASE_URL=http://127.0.0.1:8000/v1 -e OPENAI_API_KEY=none \
-e 'CORS_ALLOW_ORIGIN=http://localhost:3000;http://127.0.0.1:3000' \
-e ENABLE_PERSISTENT_CONFIG=False -v open-webui:/app/backend/data ghcr.io/open-webui/open-webui:v0.11.4
Open http://localhost:3000 and make the first account at once: it is the administrator (until it exists, whoever reaches the port first can make it), and Open WebUI asks for a login after that. Docker Desktop, and any Docker without host networking, needs the server on every interface and a key: that command is in Glyd’s README.
Measured, on an L4
On an L4 (24 GB) with no CUDA toolkit and no compiler, installed by the script from the release candidate’s own wheel, vLLM 0.30.0, Qwen3-8B unless it says otherwise. The 16 GB row is the L4 with another process holding the GPU’s memory down to what an RTX 4080 SUPER with a desktop has free (vLLM logged 14.48 GiB free at start, the figure on that card), with the plugin reading it as the GeForce Ada card it is; the 8 GB row is the same at 7.5 GiB. It was not run on an RTX 4080 SUPER, and these are not its speeds. glyd prints GB (10^9 bytes), vLLM’s log GiB.
| Free at start | glyd run chose |
vLLM logged | First start | Second |
|---|---|---|---|---|
| 23.7 GB (the whole L4) | 92% of GPU memory (21.8 GB), a 40,960-token context (the model’s own limit) | weights 11.38 GiB, KV cache 61,104 tokens | 61 to 79 s | 41 s |
| 15.7 GB (a 16 GB card with a desktop) | 62% of the L4’s memory (14.7 GB), an 11,264-token context | weights 11.39 GiB, KV cache 12,960 tokens | 78 s | 40 s |
| 8.1 GB (an 8 GB card), Qwen3-4B | refused, with the memory it needs (about 8.4 GB) and a model to try: Qwen/Qwen3-1.7B, about 5.1 GB | |||
| the same, Qwen3-1.7B | 29% (6.9 GB), a 24,576-token context | weights 2.47 GiB, KV cache 33,184 tokens | 58 s |
The first start is the first on a machine that has not run vLLM: Triton builds its launchers once, a few tens of seconds; the second is the next glyd serve. What glyd run predicts from the config was within 0.11 GiB of the weights vLLM logged and 3 to 18% under its KV cache tokens, in every run of these and of a 0.6B, a 1.7B and a 4B: it never promised a context vLLM then refused. The 18 server logs of the acceptance runs (the two cases that stop a server on purpose, Ctrl-C while loading and a killed engine, left out) have no allocator warning and no traceback. Every run with its logs, and the calibration of the constants.
The acceptance script (gpu/vllm/acceptance.sh, in v0.26.0’s tag) runs the two commands from nothing, as a user that is not root, in a container with no CUDA toolkit and no compiler: the chat page, the API, a conversation longer than the window, Open WebUI, and what a user meets around them (a program of your own at ~/.local/bin/glyd, an update, another site’s script refused, a server with a key, Ctrl-C while the model loads, uv tool uninstall glyd). It runs at a 24 GB, a 16 GB and an 8 GB card’s memory, and Glyd’s CI installs the script from each release’s tag.
Update, remove
Run the install line again to update: it installs the release the script was written for, at the package versions the acceptance run installed, and does nothing where that is what is there. uv tool uninstall glyd removes the tool and its link, and leaves what is not the tool’s: the models in the Hugging Face cache (~/.cache/huggingface, or where HF_HOME says), the last ten runs’ logs and the compiler wrapper in ~/.local/state/glyd, uv itself (~/.local/bin/uv, uv self uninstall) and its cache (~/.cache/uv, uv cache clean), and the line in your shell’s startup file if the script added one.
Advanced: vllm serve by hand
glyd run is vllm serve with its flags decided for you, so every vLLM flag works with it, and the plugin works with vllm serve as it is:
pip install "glyd[vllm]" # vLLM 0.30, and Glyd with its plugin
vllm serve Qwen/Qwen3-8B --quantization glyd # a bf16 checkpoint, packed as it loads
vllm serve ./qwen3-8b-glyd --quantization glyd # a glyd save (glyd pack, glyd.save_pretrained), as saved
The vllm extra installs vLLM 0.30, the version the plugin is tested with, beside Glyd. vLLM finds the plugin by itself; with another vLLM release, --quantization glyd stops and says why. A bf16 checkpoint is packed as it loads, a layer at a time. A Qwen3 or Llama model saved packed, by glyd pack or glyd.save_pretrained(), loads as saved.
On a 16 GB card those defaults do not leave room for a chat: vLLM sizes the context to the model’s own 40,960 tokens and takes 0.92 of the memory, which a desktop shares. By hand, for Qwen3-8B on an RTX 4080 SUPER with a desktop, with the flags and what each is for:
Needs, none of it a CUDA toolkit: an NVIDIA driver 580 or newer (the CUDA 13 build of PyTorch that vLLM 0.30 installs; the tests ran on 595), uv (it fetches Python 3.12, with its headers), a C compiler (sudo apt install build-essential: Triton builds its launchers with one, and without it vLLM stops at start with Failed to find C compiler), and disk: about 8 GB for the packages and 16 GB for the model.
uv venv --python 3.12 ~/glyd-env && source ~/glyd-env/bin/activate
uv pip install "glyd[vllm]"
VLLM_USE_FLASHINFER_SAMPLER=0 vllm serve Qwen/Qwen3-8B --quantization glyd --enforce-eager \
--max-model-len 8192 --gpu-memory-utilization 0.88 --host 127.0.0.1 \
--enable-auto-tool-choice --tool-call-parser hermes --reasoning-parser qwen3
The server answers on http://localhost:8000 once it logs Application startup complete.
VLLM_USE_FLASHINFER_SAMPLER=0: vLLM 0.30 samples top-k and top-p, which Qwen3’s own settings use, with FlashInfer, and FlashInfer builds that kernel with nvcc at the first request that samples, which vLLM’s warmup makes at start. With no CUDA toolkit the server would stop there (Could not find nvcc). This makes vLLM sample with PyTorch and Triton instead: the same tokens a second (21.25 and 21.00 for one user, greedy and with top-p, against 21.12 and 20.98 with FlashInfer’s own kernel) and the same distribution of draws.--enable-auto-tool-choice --tool-call-parser hermes: Open WebUI offers the model its built-in tools in every chat. The request carriestools, which vLLM takes astool_choiceauto, and without these flags it answers every chat with"auto" tool choice requires --enable-auto-tool-choice and --tool-call-parser to be set. hermes is Qwen3’s tool-call format.--reasoning-parser qwen3: Qwen3 thinks before it answers. With the parser the thinking is a field of its own (reasoning, which Open WebUI shows as a collapsed Thought) and the answer’scontenthas no<think>in it; without it the thinking comes in the content, tags and all.--gpu-memory-utilization 0.88is vLLM’s share of the card’s total memory, as CUDA reports it (python -c "import torch; print(torch.cuda.mem_get_info()[1] / 2**30)"), for the weights, the KV cache and the working memory: an RTX 4080 SUPER’s is 15.57 GiB (16,376 MiB is nvidia-smi’s, 15.99 GiB), so 0.88 is 13.70 GiB. The server’s CUDA context and the desktop live outside it: on that card with a desktop vLLM logged 14.48 GiB free at start. If something else holds more, vLLM stops at start withFree memory on device ... is less than desired GPU memory utilization: close that program, or lower the number.--enforce-eagerruns without torch.compile and CUDA graphs. At this budget vLLM’s defaults gave no server (−0.83 GiB left for the KV cache, measured on an L4 at 14.1 GiB); eager starts at once, and one user’s tokens a second were within 2% of a compiled server’s.--max-model-len 8192is the longest chat, in tokens. The KV cache holds 1.62 of them, and a longer chat is refused withmaximum context length is 8192 tokens.--host 127.0.0.1keeps the server on this machine. Without it vLLM listens on every interface, with no key.
Open WebUI, as above, on the server’s port 8000. To test the server alone:
curl localhost:8000/v1/models
curl localhost:8000/v1/chat/completions -H "Content-Type: application/json" -d '{
"model": "Qwen/Qwen3-8B",
"messages": [{"role": "user", "content": "What is lossless compression? Answer in one sentence. /no_think"}],
"max_tokens": 100}'
Qwen3 thinks before it answers; /no_think at the end of a message skips that.
Measured, by hand, on the L4 held to what an RTX 4080 SUPER with a desktop leaves (14.48 GiB free at start, a budget of 13.71 GiB: --gpu-memory-utilization 0.622 on the L4, where a 16,376 MiB card takes 0.88), with the plugin reading it as a GeForce Ada card; vLLM 0.30.0, Open WebUI 0.11.4. It was not run on an RTX 4080 SUPER.
- One user: 21.25 tokens a second greedy and 21.00 with top-p (the median of five answers of 256 tokens each), the first token 53 ms after the request; 161.4 and 159.5 tokens a second in all for 8 users at once.
- KV cache: 1.83 GiB, 13,280 tokens, at the 13.71 GiB budget, with the weights at 11.39 GiB.
- bf16’s weights leave no room: its server did not start with the same flags (without
--quantization glyd): it ran out of memory loading the weights, 15.26 GiB into about 15.3 GiB free. - Loading took 15 s for the weights and 39 to 46 s to a running server. While it loads, the server takes most of the free GPU memory: 669 MiB were left free at the least, with no allocator warning. No desktop ran in the test, so what one does in those seconds is not measured.
- Open WebUI: its model list showed Qwen/Qwen3-8B, a chat with its tools on completed, and it passed the model’s tool call, by
uvxand by Docker. In its page a chat showed a collapsed Thought and its answer, and a question about the time called the tool and answered with its result.
Options
The options go in --additional-config, or in the environment as GLYD_LAYOUT, GLYD_EXACT, GLYD_VERIFY, GLYD_FRACTION and GLYD_KV; anything else there is refused, with why:
vllm serve Qwen/Qwen3-8B --quantization glyd --additional-config '{"glyd": {"layout": "mma12", "verify": true}}'
layout:auto, the default, takes the layout for the GPU:mmaon Ada (L4, L40S, RTX 40) and wherever only it fits,mma12on the others (which layout on which GPU).mmais the smallest layout,mma12the 12-bit one.exact: the logits are vLLM’s bf16 ones, bit for bit (below).verify: every weight is checked against the checkpoint as it is packed, or a saved model’s against its sha256.fraction: the share of the decoder layers packed, from 0 to 1 (the default: every layer); the others stay vLLM’s own bf16 (below). A saved model takes only 1.kv:auto(the default),losslessoroff: whether vLLM’s KV cache is held in fewer bits, every value read back bit for bit (below).
vLLM’s compile cache keeps a graph for each layout and mode, so a server started with other options never runs another’s.
The lossless KV cache
From v0.27.0 vLLM’s KV cache can be held in fewer bits, with every key and value read back bit for bit (the release post). On the five models measured it holds 1.25 to 1.30 times vLLM’s tokens in the same memory, so more requests at once or longer ones, and on an A100, an L4, an H100 and a GH200 decoding was faster than with vLLM’s own cache at every batch measured. It works with any fraction, 0 included: bf16 weights with the cache held.
vllm serve Qwen/Qwen3-8B --quantization glyd # kv auto, the default
GLYD_KV=lossless vllm serve Qwen/Qwen3-8B --quantization glyd # on any GPU it supports
vllm serve Qwen/Qwen3-8B --quantization glyd --additional-config '{"glyd": {"kv": "lossless"}}'
GLYD_KV=off vllm serve Qwen/Qwen3-8B --quantization glyd # vLLM's own cache
auto, the default, holds the cache on an A100, an L4, an H100 and a GH200, by the name the GPU gives (measured on an A100 SXM4 40 GB, an L4, an H100 SXM and a GH200), and leaves vLLM’s own cache on any other GPU, with a line in the log that says so: an L40S, an RTX 40, an H200 and an A10 are not measured yet. It leaves it whereexactorverifyis on, and for a window past 40,960 tokens.losslessholds it on any GPU it supports and refuses, with why, where it cannot;offis vLLM’s own.glyd runkeeps vLLM’s own cache unlessGLYD_KVis set (GLYD_KV=auto glyd run MODEL), because of the first start below;glyd serveandvllm serve --quantization glyduseauto.- The first start sets the cache up once for the model, for the longest context the server will serve (4,096, 8,192, 16,384, 32,768 or 40,960 tokens), and keeps the result in
~/.cache/glyd/kv($GLYD_CACHE/kv). On an L4 that took 130 s for a window of 8,192 tokens (throughglyd run) and 217 s for Qwen3-8B’s whole window of 40,960 tokens (the log). A window past 40,960 tokens keeps vLLM’s own cache, with one line that says why. losslessis refused, with why, for tensor, pipeline and context parallel, speculative decoding, sliding-window, linear-attention and state-space layers, MLA, LoRA, KV connectors and offloading,--kv-cache-dtypeother than auto, head sizes other than 128, more than 8 query heads to a KV head, a model that is not bf16, another attention backend, GPUs other than Ampere, Ada and Hopper, and a model whose keys and values it cannot hold well.autoleaves vLLM’s own cache in those cases.
What is exact is the stored values: every value read back was the value written (0 of 25,683,296,256 differ in a stress run, and none of 6,329,327,616 with verify, which compares every write with a bf16 copy). Prompt logprobs are vLLM’s bit for bit: a 36,000-token prompt’s 35,999 of 35,999. A decode step’s attention is not bit-equal to vLLM’s, so greedy tokens can part from vLLM’s after some tokens (that prompt: the same for the first 19 of 64), as they do with vLLM’s own attention when one of its settings changes. With exact, which runs vLLM’s own attention for decode steps too, the 14 prompts of a prefix-hit scenario on an A100 were bit for bit, where 13 of 14 first tokens were without it (the log).
Measured, Qwen3-8B unless it says otherwise, bf16 weights, vLLM 0.30.0 in its default mode (compiled, CUDA graphs), against vLLM’s own cache. The KV cache in tokens, and decode tokens a second at the same batch (128 new tokens a request; the median of three rounds with the two alternating on the L4 and the A100, one run each on the H100 and the GH200):
| GPU | KV cache, tokens | Batch × tokens | Decode tokens a second | ||||
|---|---|---|---|---|---|---|---|
| vLLM’s own | Lossless | × | vLLM’s own | Lossless | × | ||
| L4 log | 32,208 | 41,408 | 1.286 | 1 × 1,024 | 16.6 | 16.7 | 1.007 |
| 6 × 4,096 | 77.5 | 81.6 | 1.053 | ||||
| 24 × 1,024 | 289.5 | 302.4 | 1.047 | ||||
| 1 × 8,192 | 15.5 | 15.9 | 1.023 | ||||
| 3 × 8,192 | 38.8 | 40.8 | 1.052 | ||||
| A100 SXM4 40 GB log | 136,592 | 177,392 | 1.299 | 1 × 1,024 | 75.3 | 76.7 | 1.018 |
| 1 × 8,192 | 71.4 | 73.7 | 1.032 | ||||
| H100 SXM log | 382,448 | 499,312 | 1.306 | 1 × 1,024 | 152.8 | 153.6 | 1.005 |
| 1 × 8,192 | 144.8 | 148.1 | 1.023 | ||||
| 4 × 8,192 | 499.3 | 514.5 | 1.030 | ||||
| 32 × 1,024 | 3,811.2 | 4,030.1 | 1.057 | ||||
| 32 × 8,192 | 1,733.9 | 2,075.4 | 1.197 | ||||
| 46 and 60 × 8,192 | 1,950.6 | 2,618.4 | 1.342 | ||||
| GH200 log | 482,864 | 630,112 | 1.305 | 1 × 1,024 | 176.6 | 177.2 | 1.003 |
| 1 × 8,192 | 166.4 | 168.2 | 1.011 | ||||
| 4 × 8,192 | 576.6 | 584.0 | 1.013 | ||||
| 32 × 1,024 | 4,566.2 | 4,661.5 | 1.021 | ||||
| 32 × 8,192 | 2,051.2 | 2,435.0 | 1.187 | ||||
| 58 and 76 × 8,192 | 2,395.1 | 3,126.6 | 1.305 | ||||
On the H100 and the GH200 the last row is each cache at the largest batch it holds: 46 requests of 8,192 tokens with vLLM’s own, 60 with the lossless cache; the GH200’s, 58 and 76.
- The KV cache on five models, an L4, eager,
--max-model-len 8704: Qwen3-4B 1.2982, Qwen3-8B 1.2834, Qwen2.5-7B 1.2773, Mistral-7B-v0.3 1.2743, Llama-3.1-8B 1.2526 times vLLM’s tokens (log). - Serving, 1,024 tokens in and 256 out, saturated, every pass with prompts of its own, the lossless cache alone served 1.23 times vLLM’s requests a second on an L4 (0.93 against 0.76), 1.29 on an A100 (7.62 against 5.92), 1.05 on an H100 SXM (21.68 against 20.58) and 1.06 on a GH200 (22.47 against 21.18). On the L4 with Glyd’s weights as well, against bf16 and vLLM’s own cache, it served 1.59 times the requests a second (1.20 against 0.76) with 2.64 times the KV tokens (71,248 against 27,024), its first token after 0.58 times the time and each token 1.55 times as long (L4, A100, H100, GH200). The prefix cache’s hit rate in the servers’ logs was 0.0% on the L4 and the GH200, at most 0.9% on the A100 and 1.4% on the H100.
- With long prompts, an H100 SXM trades speed for room: 8,192 tokens in and 256 out, saturated, it served 1.05 times vLLM’s requests a second (2.80 against 2.67) and its first token came after 0.89 times the time (8,196 ms against 9,199), but each generated token took 1.12 times as long (49.2 ms against 43.9). The KV cache held 494,256 tokens against 377,024 (1.31 times) and the model’s memory was 1.07 times (16.34 against 15.27 GiB) (log). On a GH200 the same run served 1.12 times the requests a second (3.12 against 2.78), with its first token after 0.95 times the time and each token 0.97 times as long (log). An H200 is not measured.
- Prompts of a model’s rare tokens (tokens whose embeddings were never trained: Llama-3.1-8B has 289, Mistral-7B-v0.3 902, the Qwen models none) are the cache’s known limit. 40 of them of 1,400 tokens at once with 8 ordinary prompts: all 40 served on every one of five models, none ended with an error; the flood ran at 0.88 times (Llama-3.1-8B) and 0.77 times (Mistral-7B-v0.3) vLLM’s prompt tokens a second, and at 1.00 to 1.01 times on the Qwen models. The ordinary prompts sent with it were not held behind it: 8.1 s on Mistral-7B-v0.3 and 10.5 s on Llama-3.1-8B, where vLLM’s own took 19.0 s and 19.8 s.
- The checks against vLLM’s own pass on all five models (13 of 13 on Qwen3-4B and Qwen2.5-7B, 21 of 21 on Qwen3-8B, Mistral-7B-v0.3 and Llama-3.1-8B); Llama-3.1-8B’s full run passes 34 of 35, the one miss being the prefix-hit check, where 13 of 14 first tokens were vLLM’s bit for bit. On the H100 and the GH200, the kernels’ 81 of 81 checks and the prefix-caching, chunked-prefill and
exactchecks, 12 of 12, pass.
A fraction of the layers
A packed layer saves memory; a layer left as it is saves nothing. fraction says how many layers are packed, spread evenly over the model’s depth (0.5 packs every second layer), a layer’s Linears together and a mixture of experts’ experts with them. 0 is vLLM’s own bf16, with Glyd’s library not loaded; 1 is every layer.
vllm serve Qwen/Qwen3-32B --quantization glyd --additional-config '{"glyd": {"fraction": 0.5}}'
Measured on an H100 SXM (80 GB) with Qwen3-32B, every request sent at once (192 prompts, 1,024 tokens in and 256 out), each server started cold with --max-model-len 4096 --max-num-seqs 128 and the same --gpu-memory-utilization (0.9193), one run each, against bf16 in the same job. Glyd’s layout there is the 12-bit one:
| Layers packed | Weights, GiB | KV cache | Requests a second | First token, ms | Each token, ms | |||
|---|---|---|---|---|---|---|---|---|
| Tokens | × | Number | × | |||||
| bf16 | 61.03 | 32,320 | 1.00 | 2.51 | 1.00 | 32,958 | 36.6 | |
| fraction 0 | 0 of 64 | 61.03 | 32,320 | 1.00 | 2.51 | 1.00 | 32,982 | 36.6 |
| fraction 0.25 | 16 of 64 | 58.20 | 43,872 | 1.36 | 2.85 | 1.14 | 27,116 | 41.7 |
| fraction 0.5 | 32 of 64 | 54.89 | 57,392 | 1.78 | 3.54 | 1.41 | 23,198 | 46.2 |
| fraction 0.75 | 48 of 64 | 51.59 | 70,912 | 2.19 | 3.50 | 1.40 | 20,883 | 52.8 |
| fraction 1 | 64 of 64 | 48.27 | 84,528 | 2.62 | 3.91 | 1.56 | 20,014 | 56.8 |
- On an 80 GB H100, Qwen3-32B’s bf16 weights leave room for only 32,320 tokens of KV cache, and every packed fraction served more requests a second than bf16: 1.14 times at 0.25, 1.41 at 0.5, 1.40 at 0.75 and 1.56 with every layer packed. bf16 ran at most 25 requests at a time, with up to 175 waiting; with every layer packed, at most 81.
- The first token came sooner at every packed fraction (0.61 times bf16’s 33.0 s with every layer packed) and each token later (1.14 to 1.55 times). At fraction 1 the GPU drew a median of 698 W of its 700 W limit.
- Fraction 0 was bf16: 2.506 against 2.507 requests a second, each token 36.6 ms in both.
One GPU, one model and one load: the first token at 1 request a second, a fraction on a GH200, and other models were not measured. The run and its logs.
Exact mode
By default the products round a little differently from bf16’s, as any two GPU kernels do, so a late token can differ from bf16’s. vLLM’s own bf16 differs the same way with and without CUDA graphs. With exact, the products are vLLM’s bf16 ones, and the logits are bf16’s bit for bit. That holds with --enforce-eager, or compiled in inductor’s deterministic mode, the one mode in which vLLM’s compiled bf16 is itself the same from one run to the next:
vllm serve Qwen/Qwen3-8B --quantization glyd --additional-config '{"glyd": {"exact": true}}' \
--compilation-config '{"inductor_compile_config": {"deterministic": true, "combo_kernels": true, "benchmark_combo_kernel": false}}'
Asked for exact with neither, the server stops at start and says why. The deterministic mode cost nothing measurable on an L4. For a model whose Linears have biases, such as Qwen2.5, exact runs eager: compiled bf16 adds those biases in a step of their own, so the bits can differ there.
Measured
vllm bench serve, bf16 against Glyd at the same --gpu-memory-utilization 0.9: servers warm, 1,024 tokens in and 256 out, vLLM 0.30.0. One GPU each; Qwen3-30B-A3B is a mixture of experts. The A100’s, the GH200’s and the H100’s rows were measured with each pass’s prompts new to the server; the L4’s and the A10’s saturated passes served none from vLLM’s prefix cache. Low load is 1 request a second, 0.25 on the L4; saturated is every request at once. Glyd’s layout is the one auto picks: mma on the L4, mma12 on the others. For the times, lower is better. 2× RTX A6000: being re-measured after a benchmark fix.
| GPU, model | KV cache, × | Requests a second saturated, × | Low load, % | Saturated, % | ||
|---|---|---|---|---|---|---|
| First token | Each token | First token | Each token | |||
| L4, Qwen3-8B log | 1.89 | 1.33 | +16 | −21 | −25 | +42 |
| A10, Qwen3-8B log | 1.73 | 1.31 | +16 | −21 | −25 | +30 |
| A100 40 GB, Qwen3-8B log | 1.14 | 0.98 | +19 | −10 | +9 | +17 |
| A100 40 GB, Qwen3-14B log | 1.77 | 1.28 | +19 | −13 | −21 | +33 |
| GH200, Qwen3-8B log | 1.04 | 0.93 | +6 | +1 | +5 | +9 |
| GH200, Qwen3-32B log | 1.66 | 0.89 | +27 | −6 | −24 | +58 |
| H100 SXM, Qwen3-30B-A3B log | 2.11 | 0.84 | +35 | +5 | −37 | +126 |
- More requests at once everywhere: 1.04 to 2.11 times bf16’s KV cache.
- More requests a second saturated on the L4, A10 and, with Qwen3-14B, the A100: 1.28 to 1.33 times, with the first token 21 to 25% sooner; with Qwen3-8B on the A100, where bf16 has room to spare, 0.98 times. Qwen3-32B on an 80 GB H100 SXM, where bf16 is short of KV cache, served 1.56 times with every layer packed (above; a bench of its own, not the table’s).
- Each token at low load: 10 to 21% sooner on the L4, A10 and A100, and 6% sooner with Qwen3-32B on the GH200; 1% later with Qwen3-8B on the GH200 and 5% later with Qwen3-30B-A3B on the H100 SXM.
- Slower: the first token at low load, 6 to 35% later. On the GH200 saturated, 0.89 to 0.93 times bf16’s requests a second; with Qwen3-30B-A3B on the H100 SXM, 0.84 times, although bf16 filled its KV cache (117 requests running, 139 waiting; Glyd ran up to 222); its experts’ products at those batches are the next work. Each token at saturation, 30 to 42% later on the L4 and A10, 17% and 33% on the A100, 9 and 58% on the GH200 and 126% with Qwen3-30B-A3B on the H100 SXM.
Every weight is served bit for bit, and the tokens differ from bf16’s about as much as vLLM’s own bf16 differs between its modes. Checked against vLLM’s own bf16 (check_vllm.py): Qwen3-8B on an L4, an A10, an A100, a GH200 and an H100 SXM (the H100 SXM’s 15 checks all passed), and on an L4 five more models, a mixture of experts and a model with biases among them. Every rate and percentile, those checks and how to run them, as of v0.26.0: gpu/vllm.
Speculative decoding
vLLM’s speculative decoding runs on Glyd’s packed weights too. On an L4, Qwen3-8B, one user, greedy, output tokens a second on prompts that edit a given text or code, and on chat (log):
| Output tokens/s | Edit | Chat |
|---|---|---|
| bf16 | 16.6 | 16.7 |
| bf16, EAGLE-3 draft | 43.3 | 30.3 |
| Glyd | 21.2 | 21.5 |
| Glyd, n-gram lookup | 35.5 | 22.1 |
| Glyd, EAGLE-3 draft | 55.0 | 38.9 |
The EAGLE-3 draft is RedHatAI/Qwen3-8B-speculator.eagle3. With bf16 it fit the L4 only with --gpu-memory-utilization 0.95. With speculation, exact eager gives bf16 eager’s tokens. Speculative greedy tokens can differ from plain decoding’s, in vLLM’s bf16 too: a verify step multiplies several tokens at once, and the GPU kernels round by that shape. Under VLLM_BATCH_INVARIANT=1 they are plain decoding’s, bf16’s and exact’s alike.
Mixtures of experts
A mixture of experts’ experts are packed too. Measured on granite-3.1-3b-a800m-instruct on an L4: every expert bit for bit, and exact bf16’s bits (log). Qwen3-30B-A3B over two RTX A6000s with tensor parallelism: every expert bit for bit, exact bf16’s bits; its serving numbers are being re-measured after a benchmark fix. Qwen3-30B-A3B on one H100 SXM: 2.11 times bf16’s KV cache and 0.84 times its requests a second saturated, in the table above.
Not supported yet
- LoRA adapters, dual-batch overlap, weight offloading and sleep mode: the server stops at start and says so.
- A model saved packed, served over several GPUs, with a mixture of experts’ packs, or of a family other than Qwen3’s and Llama’s: serve its bf16 checkpoint.
exactcompiled for a model whose Linears have biases, and Glyd’s own kernels underVLLM_BATCH_INVARIANT: the server stops at start and says why.- Faster than bf16 everywhere: the known gaps.