MiniMax-M3
MiniMaxAI/MiniMax-M3 on Hugging Face
Weights in GPU memory
854 to 575 GB
About 33% less, worked out from the parameter count; run queued est.8×H100 nodes it takes
2 to 1
Weights and an 8K-token KV cache on 80 GB H100sH100 80 GB GPUs it takes
11 to 7
Weights, an 8K-token KV cache and 1.5 GB a GPU for the runtimeWhich single GPU it fits
Worked out from the parameter count; not measured yet.
| Memory | GPUs | bf16 | Glyd |
|---|---|---|---|
| 16 GB | RTX 4080, RTX 5080 | No | No |
| 24 GB | RTX 4090, RTX 3090, A10 | No | No |
| 32 GB | RTX 5090 | No | No |
| 48 GB | RTX A6000, L40S, RTX 6000 Ada | No | No |
| 80 GB | H100, A100 80 GB | No | No |
| 96 GB | RTX PRO 6000, GH200 | No | No |
| 141 GB | H200 | No | No |
Measured so far
21 open models measured, with quality wherever bf16 fits one GPU.