4x NVIDIA RTX 3090
Desktop1
Run records
1
Models covered
1
Frameworks covered
L0 Self-reported × 1
Evidence mix
Typical measured performance
Published benchmark averages per model × framework × quantization on this hardware.
- llama.cpp · UD-IQ2_XXS Decode 35.68 · Prefill 420.73 ·
Key specs
- VRAM
- 96 GB
Run records
1 configurations
| Model | Framework | Quantization | Decode | Samples |
|---|---|---|---|---|
| DeepSeek-V4-Flash | llama.cpp | UD-IQ2_XXS | 35.7 tok/s | 1 |