Release · Oct 1, 2026 · Glyd
Glyd v0.26.0: one command to a local chat, serving with vLLM
Glyd v0.26.0 installs with one command and starts a model with glyd run; vLLM serves it packed, now measured on an H100 SXM, and long prompts run faster.
Glyd v0.26.0 starts a model with one command. The script installs Glyd with vLLM, and glyd run downloads a model, packs its weights as they load, and opens a chat in the terminal and at http://localhost:8000:
curl -LsSf https://getglyd.com/install.sh | sh
glyd run Qwen/Qwen3-8B
It needs Linux, an NVIDIA GPU of the Ampere generation or newer, and its driver, 580 or newer; no CUDA toolkit and no sudo. On a Mac, or with no NVIDIA GPU, the script installs the compression program instead and says what glyd run needs. v0.26.0 also serves models in vLLM with --quantization glyd, measured now on an H100 SXM too, and runs long prompts faster on an A100 SXM, a GH200 and an H100 SXM.
One command to a chat
The script installs uv where there is none, then Glyd with vLLM 0.30 and PyTorch as one isolated tool, on a Python 3.12 that uv fetches: about 8 GB, all under your home directory, with no sudo, no virtual environment and no pip. It ends with glyd doctor, which says what the machine can run. glyd run MODEL then:
- checks the GPU, its driver and the model, and whether the model with Glyd fits the memory free now (if not, it says how much it needs, and the largest model of its family that fits);
- downloads the model, and works the settings out from the GPU: vLLM’s share of the memory and the longest context its KV cache holds;
- starts vLLM, and chats in the terminal (
/bye,/clear,/think) and on a page athttp://localhost:8000.
glyd serve MODEL leaves the server up as an OpenAI API for other programs, and Open WebUI runs in front of it with its login on. The commands, the settings and what the script does to a machine: local chat, like Ollama.
Measured on an L4 held to what an RTX 4080 SUPER with a desktop leaves free of its 16 GB, 14.48 GiB (it was not run on that card): Qwen3-8B with Glyd took 11.39 GiB of weights, and glyd run chose an 11,264-token context with a KV cache of 12,960 tokens; by hand, with the flags of the docs for that card (an 8,192-token context), the cache held 13,280 tokens and one user got 21.25 tokens a second. bf16’s server did not start with the same flags: it ran out of memory loading the weights. The loads logged no allocator warning, and no CUDA toolkit was installed (the runs, the by-hand log).
Serving with vLLM
Dense models and mixtures of experts, on one GPU or over two, with speculative decoding; exact gives vLLM’s bf16 logits bit for bit. vllm bench serve, bf16 against Glyd at the same memory setting, 1,024 tokens in and 256 out:
| GPU, model | KV cache, × | Requests a second saturated, × |
|---|---|---|
| L4, Qwen3-8B log | 1.89 | 1.33 |
| A10, Qwen3-8B log | 1.73 | 1.31 |
| A100 40 GB, Qwen3-8B log | 1.14 | 0.98 |
| A100 40 GB, Qwen3-14B log | 1.77 | 1.28 |
| GH200, Qwen3-8B log | 1.04 | 0.93 |
| GH200, Qwen3-32B log | 1.66 | 0.89 |
| H100 SXM, Qwen3-30B-A3B log | 2.11 | 0.84 |
More requests a second saturated on the L4, A10 and, with Qwen3-14B, the A100, as many with Qwen3-8B on the A100, fewer on the GH200 and for the mixture of experts, Qwen3-30B-A3B, on an H100 SXM (the known gaps). At low load each token comes 10 to 21% sooner on the L4, A10 and A100; the first token comes 6 to 35% later on every GPU measured. The A100’s, the GH200’s and the H100’s rows are as measured again in v0.27.0, with each pass’s prompts new to the server; the two RTX A6000s’ row is out until it is re-measured after a benchmark fix.
One user, speculative decoding, Qwen3-8B on an L4, output tokens a second (the log):
| Output tokens/s | Edit a text | Chat |
|---|---|---|
| bf16 | 16.6 | 16.7 |
| bf16, EAGLE-3 draft | 43.3 | 30.3 |
| Glyd, EAGLE-3 draft | 55.0 | 38.9 |
With EAGLE-3, Glyd made 3.3 times bf16’s tokens a second on the edit mix and 2.3 times on chat, and 1.27 to 1.28 times bf16’s with the same draft; bf16 with the draft fit the L4 only with --gpu-memory-utilization 0.95 --max-num-batched-tokens 2048.
The options, exact mode and what is not supported yet: serving with vLLM.
An 80 GB H100 SXM and a 32B model
Qwen3-32B’s bf16 weights, 61.03 GiB, leave an 80 GB H100 SXM room for 32,320 tokens of KV cache. Glyd’s fraction option packs a share of the layers and leaves the others bf16. Every request sent at once (192 prompts, 1,024 tokens in and 256 out), each server started cold, against bf16 in the same job:
| Layers packed | KV cache, × | Requests a second, × | First token, % | Each token, % | |
|---|---|---|---|---|---|
| fraction 0.25 | 16 of 64 | 1.36 | 1.14 | −18 | +14 |
| fraction 0.5 | 32 of 64 | 1.78 | 1.41 | −30 | +26 |
| fraction 0.75 | 48 of 64 | 2.19 | 1.40 | −37 | +44 |
| fraction 1 | 64 of 64 | 2.62 | 1.56 | −39 | +55 |
Every packed fraction served more requests a second than bf16, which ran at most 25 requests at a time with up to 175 waiting: 1.56 times with every layer packed, which held 2.62 times the KV cache and ran up to 81 at a time. The first token came sooner at every packed fraction and each token later. Fraction 0 is bf16: 2.506 against 2.507 requests a second. One run each on one GPU, at full load; the first token at low load was not measured here (the log). The same GPU with the mixture of experts Qwen3-30B-A3B, in the table above: 2.11 times the KV cache and 0.95 times the requests a second, where bf16 filled its KV cache.
Faster long prompts on an A100 SXM, a GH200 and an H100 SXM
Long prompts on these GPUs run faster. A forward pass where the faster route runs, its time over v0.25.1’s (— where it does not):
| GPU, model | Prompt, tokens | ||||
|---|---|---|---|---|---|
| 769 | 1,024 | 2,048 | 4,096 | 8,192 | |
| A100 SXM, Qwen3-8B | 0.899 | 0.876 | 0.959 | 0.971 | — |
| A100 SXM, Qwen3-14B | 0.845 | 0.864 | 0.916 | 0.947 | 0.968 |
| GH200, Qwen3-32B | — | — | 0.909 | 0.940 | 0.952 |
| H100 SXM, Qwen3-14B | — | — | 0.893 | 0.937 | 0.938 |
The logs, the H100 SXM’s in its h100-sxm-measure. Where it runs, as measured: the CUDA library.
How fast it responds
A new benchmark times what a user of generate() waits for: the first token, tokens a second, and a chat and a long document to their last token, Glyd’s default against bf16 compiled the same way, on an L4, an A10, an A100 SXM and a GH200. Glyd’s default made 1.06 to 1.31 times bf16’s tokens a second at one sequence and 0.94 to 1.33 times at 8 and 32, and Qwen3-14B runs on an A10 where bf16 does not fit (how fast it responds).
The C API
The CUDA library’s C API is version 7: the calls for the faster long-prompt route (the CUDA library).
Where Glyd is still slower than bf16: the known gaps. Every change, with its numbers: the changelog and the release notes.