Benchmarks
Measured on real GPUs, bf16 and Glyd in the same run. Every number links to the run's raw log, and every run can be repeated from the repository.
GPU time per generated token
Qwen2.5-7B-Instruct. Lower is better. Each GPU runs the layout Glyd picks for it.
1 sequence a step
8 sequences a step
32 sequences a step
64 sequences a step
Against DFloat11 on a 16 GB GPU
Qwen3-8B on an RTX 4080 SUPER, where its 16.38 GB of bf16 does not fit. Tokens per second, higher is better.
| Sequences a step | DFloat11 | Glyd |
|---|---|---|
| 1 sequence 3.4× | 13.8 | 47.2 |
| 8 sequences 3.4× | 104.7 | 360.2 |
| 32 sequences 3.3× | 387.5 | 1268.6 |
| 64 sequences 2.5× | 746.2 | 1897.2 |
Both hold the weights in about 11.2 GB. DFloat11 decompresses each block to bf16 before running it; Glyd decodes inside the product.
One 48 GB GPU instead of two
Qwen3-32B on RTX A6000s: bf16 across two GPUs, Glyd on one. Tokens per second, higher is better.
| Sequences a step | bf16, 2 GPUs | Glyd, 1 GPU |
|---|---|---|
| 1 sequence +24% | 9.5 | 11.8 |
| 8 sequences +28% | 74 | 95 |
| 32 sequences +5% | 276 | 290 |
Qwen2.5-72B the same way: four 48 GB GPUs in bf16, three with Glyd, 4.5 → 6.4 tokens/s at one sequence.
The matrix multiply against ZipServ and cuBLAS
Qwen3-8B, layer 18's matrices, RTX 4080 SUPER, microseconds a call with the L2 cache flushed (as ZipServ times itself). Lower is better; Glyd's time in blue where it is the fastest of the three.
| Matrix | Bits a weight, ZipServ / Glyd | 1 token | 16 tokens | 32 tokens | 64 tokens |
|---|---|---|---|---|---|
| q_proj, o_proj (4096 × 4096) | 11.35 / 10.81 | 73 / 57 / 56 | 73 / 57 / 58 | 74 / 58 / 62 | 77 / 69 / 73 |
| k_proj (1024 × 4096) | 11.43 / 10.83 | 23 / 22 / 19 | 27 / 23 / 21 | 26 / 23 / 24 | 28 / 25 / 30 |
| gate_proj (12288 × 4096) | 11.35 / 10.76 | 175 / 151 / 144 | 204 / 152 / 148 | 235 / 155 / 153 | 225 / 162 / 172 |
| down_proj (4096 × 12288) | 11.35 / 10.75 | 176 / 153 / 147 | 208 / 154 / 150 | 228 / 155 / 153 | 214 / 187 / 181 |
How we measure
- Speed
- GPU time of a generated token apart from the prompt: 17 steps less 1, over 16. Greedy decoding.
- Like a server
- q, k, v as one product and gate, up as another, for bf16 and Glyd alike, as vLLM runs them.
- Quality
- Perplexity on enwik8 and MMLU questions, bf16 against Glyd in the same run.
- Exactness
- Every matrix packed, unpacked and compared bit for bit.
- Hardware
- Own RTX 4080 SUPER; Lambda Cloud A10, A100, H100 and RTX A6000. Driver, CUDA and PyTorch versions in each log.
Known gaps
- Many sequences on A100 and H10032 to 64 sequences a step take 6 to 17% more GPU time than bf16 today.
- FP8 and 4-bit modelsFP8 weights can shrink 16 to 18% (measured), but the GPU kernels for them are not built yet. 4-bit weights have little left to take.
- Long prompts on HopperPrompt processing on H100 does not use the fused kernels yet.
- BlackwellNot measured yet on RTX 50-series, RTX PRO 6000 or B200.