DeepSeek-V4-Flash · UD-IQ2_XXS
L0 自报2x NVIDIA RTX 5090 · 64GB · llama.cpp
GitHub
今天
19.39 tok/s
Decode
692.41 tok/s
Prefill
0.12 s
TTFT
51.17 GB
VRAM
配置(最近一次实测):
Ubuntu 24.04.4 LTS CUDA 13.3 UD-IQ2_XXS llama.cpp b10236
来源:GitHub · 1 次实测 · 1 个独立来源
2x NVIDIA RTX 5090 · 64GB · llama.cpp
GitHub
今天
19.39 tok/s
Decode
692.41 tok/s
Prefill
0.12 s
TTFT
51.17 GB
VRAM
配置(最近一次实测):
来源:GitHub · 1 次实测 · 1 个独立来源
4x NVIDIA RTX 3090 · 96GB · llama.cpp
GitHub
今天
35.68 tok/s
Decode
420.73 tok/s
Prefill
0.11 s
TTFT
87.16 GB
VRAM
配置(最近一次实测):
来源:GitHub · 1 次实测 · 1 个独立来源
共 2 条结果