DeepSeek-V4-Flash · UD-IQ2_XXS
L0 Self-reported2x NVIDIA RTX 5090 · 64GB · llama.cpp
GitHub
Today
19.39 tok/s
Decode
692.41 tok/s
Prefill
0.12 s
TTFT
51.17 GB
VRAM
Config (latest run):
Ubuntu 24.04.4 LTS CUDA 13.3 UD-IQ2_XXS llama.cpp b10236
Sources: GitHub · 1 runs · 1 independent sources