Glyd

Benchmarks

Measured on real GPUs, bf16 and Glyd in the same run. Every number links to the run's raw log, and every run can be repeated from the repository.

Updated Sep 27, 2026RTX 4080 SUPER · RTX A6000 · A10 · A100 · H100

GPU time per generated token

Qwen2.5-7B-Instruct. Lower is better. Each GPU runs the layout Glyd picks for it.

1 sequence a step

RTX 4080 SUPER, 16 GBTiered layout, 10.80 bits
bf16: 21.97 ms
Glyd: 16.52 ms
25% less time
A10, 24 GB12-bit layout
bf16: 33.71 ms
Glyd: 24.40 ms
28% less time
A100, 40 GB12-bit layout
bf16: 14.18 ms
Glyd: 11.84 ms
17% less time
H100 SXM, 80 GB12-bit layout
bf16: 6.90 ms
Glyd: 6.55 ms
5% less time

8 sequences a step

RTX 4080 SUPER, 16 GBTiered layout, 10.80 bits
bf16: 22.85 ms
Glyd: 17.37 ms
24% less time
A10, 24 GB12-bit layout
bf16: 34.77 ms
Glyd: 25.76 ms
26% less time
A100, 40 GB12-bit layout
bf16: 14.87 ms
Glyd: 13.77 ms
7% less time
H100 SXM, 80 GB12-bit layout
bf16: 7.50 ms
Glyd: 7.24 ms
3% less time

32 sequences a step

RTX 4080 SUPER, 16 GBTiered layout, 10.80 bits
bf16: 26.19 ms
Glyd: 18.80 ms
28% less time
A10, 24 GB12-bit layout
bf16: 35.50 ms
Glyd: 29.03 ms
18% less time
A100, 40 GB12-bit layout
bf16: 16.18 ms
Glyd: 17.16 ms
6% more time
H100 SXM, 80 GB12-bit layout
bf16: 8.02 ms
Glyd: 8.69 ms
8% more time

64 sequences a step

RTX 4080 SUPER, 16 GBTiered layout, 10.80 bits
bf16: 27.67 ms
Glyd: 25.14 ms
9% less time
A10, 24 GB12-bit layout
bf16: 37.98 ms
Glyd: 33.05 ms
13% less time
A100, 40 GB12-bit layout
bf16: 17.80 ms
Glyd: 19.95 ms
12% more time
H100 SXM, 80 GB12-bit layout
bf16: 8.55 ms
Glyd: 9.99 ms
17% more time
bf16GlydA step after the prompt, q, k, v and gate, up merged as vLLM runs them. Bars scaled within each row.

Against DFloat11 on a 16 GB GPU

Qwen3-8B on an RTX 4080 SUPER, where its 16.38 GB of bf16 does not fit. Tokens per second, higher is better.

Sequences a stepDFloat11Glyd
1 sequence 3.4×13.847.2
8 sequences 3.4×104.7360.2
32 sequences 3.3×387.51268.6
64 sequences 2.5×746.21897.2

Both hold the weights in about 11.2 GB. DFloat11 decompresses each block to bf16 before running it; Glyd decodes inside the product.

One 48 GB GPU instead of two

Qwen3-32B on RTX A6000s: bf16 across two GPUs, Glyd on one. Tokens per second, higher is better.

Sequences a stepbf16, 2 GPUsGlyd, 1 GPU
1 sequence +24%9.511.8
8 sequences +28%7495
32 sequences +5%276290

Qwen2.5-72B the same way: four 48 GB GPUs in bf16, three with Glyd, 4.5 → 6.4 tokens/s at one sequence.

The matrix multiply against ZipServ and cuBLAS

Qwen3-8B, layer 18's matrices, RTX 4080 SUPER, microseconds a call with the L2 cache flushed (as ZipServ times itself). Lower is better; Glyd's time in blue where it is the fastest of the three.

MatrixBits a weight, ZipServ / Glyd1 token16 tokens32 tokens64 tokens
q_proj, o_proj (4096 × 4096)11.35 / 10.8173 / 57 / 5673 / 57 / 5874 / 58 / 6277 / 69 / 73
k_proj (1024 × 4096)11.43 / 10.8323 / 22 / 1927 / 23 / 2126 / 23 / 2428 / 25 / 30
gate_proj (12288 × 4096)11.35 / 10.76175 / 151 / 144204 / 152 / 148235 / 155 / 153225 / 162 / 172
down_proj (4096 × 12288)11.35 / 10.75176 / 153 / 147208 / 154 / 150228 / 155 / 153214 / 187 / 181

Each cell: cuBLAS bf16 / ZipServ / Glyd. Glyd's tiered layout at 1 to 32 tokens, its 12-bit layout at 64.

How we measure

Speed
GPU time of a generated token apart from the prompt: 17 steps less 1, over 16. Greedy decoding.
Like a server
q, k, v as one product and gate, up as another, for bf16 and Glyd alike, as vLLM runs them.
Quality
Perplexity on enwik8 and MMLU questions, bf16 against Glyd in the same run.
Exactness
Every matrix packed, unpacked and compared bit for bit.
Hardware
Own RTX 4080 SUPER; Lambda Cloud A10, A100, H100 and RTX A6000. Driver, CUDA and PyTorch versions in each log.
Every run's raw log

Known gaps

  • Many sequences on A100 and H10032 to 64 sequences a step take 6 to 17% more GPU time than bf16 today.
  • FP8 and 4-bit modelsFP8 weights can shrink 16 to 18% (measured), but the GPU kernels for them are not built yet. 4-bit weights have little left to take.
  • Long prompts on HopperPrompt processing on H100 does not use the fused kernels yet.
  • BlackwellNot measured yet on RTX 50-series, RTX PRO 6000 or B200.