Explainers on how Glyd makes AI model weights smaller without changing a bit of them.
A bf16 weight spends 8 of its 16 bits on the exponent, but in a trained model those 8 bits carry about 2.6 bits of information. Why, and what it buys.
Also as RSS.