Mistral Medium 3.5 128B
mistralai/Mistral-Medium-3.5-128B on Hugging Face
Weights in GPU memory
134 to 130 GB
FP8 weights counted as they are: Glyd’s GPU kernels for them are not built yet est.8×H100 nodes it takes
1 to 1
Weights and an 8K-token KV cache on 80 GB H100sH100 80 GB GPUs it takes
2 to 2
Weights, an 8K-token KV cache and 1.5 GB a GPU for the runtimeWhich single GPU it fits
FP8 weights counted as they are; not measured yet.
| Memory | GPUs | Released | Glyd |
|---|---|---|---|
| 16 GB | RTX 4080, RTX 5080 | No | No |
| 24 GB | RTX 4090, RTX 3090, A10 | No | No |
| 32 GB | RTX 5090 | No | No |
| 48 GB | RTX A6000, L40S, RTX 6000 Ada | No | No |
| 80 GB | H100, A100 80 GB | No | No |
| 96 GB | RTX PRO 6000, GH200 | No | No |
| 141 GB | H200Room for the KV cache: about 43K tokens as released, 54K with Glyd | Fits | Fits |
Measured so far
21 open models measured, with quality wherever bf16 fits one GPU.