Reports
Measured runs of Glyd against bf16 and the other lossless formats, each with its method, the command to repeat it and a link to the raw logs.
- Report ·
Against DFloat11 and ZipServ on an RTX 4080 SUPER
Qwen3-8B on a 16 GB RTX 4080 SUPER: Glyd generates 2.5 to 3.4 times DFloat11's tokens a second at the same size, and its kernels keep pace with ZipServ's.
- Report ·
Qwen3-32B on one 48 GB GPU, bit for bit
Glyd holds Qwen3-32B's weights in 44.5 GB instead of 65.5, so it runs on one 48 GB RTX A6000 instead of two, every weight exact and faster.
- Report ·
Ten popular open models: every matrix 32 to 33% smaller, bit for bit
Every Linear layer's matrix of ten popular open models, from SmolLM3 3B to Llama 3.3 70B, packed and unpacked bit for bit: 31.9 to 33.0% smaller.