Glyd

Mistral Small 4 119B

  • Mistral AI
  • 119.4B parameters
  • Mixture of experts, 6.5B active
  • Apache-2.0
  • Released in FP8
  • Queued for measurement

mistralai/Mistral-Small-4-119B-2603 on Hugging Face

Weights in GPU memory
121 to 120 GB
FP8 weights counted as they are: Glyd’s GPU kernels for them are not built yet est.
Smallest single GPU
141 GB to 141 GB
Weights, an 8K-token KV cache and 1.5 GB for the runtime
H100 80 GB GPUs it takes
2 to 2
Weights, an 8K-token KV cache and 1.5 GB a GPU for the runtime

Which single GPU it fits

FP8 weights counted as they are; not measured yet.

MemoryGPUsReleasedGlyd
16 GBRTX 4080, RTX 5080NoNo
24 GBRTX 4090, RTX 3090, A10NoNo
32 GBRTX 5090NoNo
48 GBRTX A6000, L40S, RTX 6000 AdaNoNo
80 GBH100, A100 80 GBNoNo
96 GBRTX PRO 6000, GH200NoNo
141 GBH200Room for the KV cache: about 1.2M tokens as released, 1.3M with GlydFitsFits

Measured so far

21 open models measured, with quality wherever bf16 fits one GPU.