Glyd

DeepSeek-V4 Pro

  • DeepSeek
  • Mixture of experts, 49B active
  • MIT
  • Released in 4-bit
  • Queued for measurement

deepseek-ai/DeepSeek-V4-Pro on Hugging Face

Weights in GPU memory
865 to 863 GB
4-bit weights counted as they are: Glyd’s GPU kernels for them are not built yet est.
8×H100 nodes it takes
2 to 2
Weights and an 8K-token KV cache on 80 GB H100s
H100 80 GB GPUs it takes
11 to 11
Weights, an 8K-token KV cache and 1.5 GB a GPU for the runtime

Which single GPU it fits

4-bit weights counted as they are; not measured yet.

MemoryGPUsReleasedGlyd
16 GBRTX 4080, RTX 5080NoNo
24 GBRTX 4090, RTX 3090, A10NoNo
32 GBRTX 5090NoNo
48 GBRTX A6000, L40S, RTX 6000 AdaNoNo
80 GBH100, A100 80 GBNoNo
96 GBRTX PRO 6000, GH200NoNo
141 GBH200NoNo

Measured so far

21 open models measured, with quality wherever bf16 fits one GPU.