Qwen3-32B
L0 Self-reportedQ4_K_M · llama.cpp b4567
- NVIDIA RTX 4090
- 24GB
- Ubuntu 24.04
- CUDA 12.4
18tok/s
Decode
210.5tok/s
Prefill
19.5GB
VRAM
- NVIDIA
- llama.cpp
- Self-reported
Query the model, framework, and quantization combos your hardware can run — based on real benchmark runs.
Beyond speed: how reliable each data point is, whether it's directly comparable, and where the evidence comes from.
11
Records
Real benchmark runs
4
Hardware
Mainstream GPUs / CPUs
3
Models
LLMs & specialized models
3
Frameworks
llama.cpp / vLLM & more
The latest model run results from real collected data.
Q4_K_M · llama.cpp b4567
18tok/s
Decode
210.5tok/s
Prefill
19.5GB
VRAM
Q4_K_M · llama.cpp b4602
18.6tok/s
Decode
218tok/s
Prefill
19.4GB
VRAM
Q4_K_M · llama.cpp b4602
18.6tok/s
Decode
222.3tok/s
Prefill
19.6GB
VRAM
Q4_K_M · Ollama 0.5.7
17.4tok/s
Decode
198.2tok/s
Prefill
19.8GB
VRAM
A structured evidence system tells you how much each data point can be trusted.
Published by the community or individuals, not independently verified.
Community reproductions — small-scale independent verification with reproduction counts.
Multi-framework benchmark runs on the same hardware + model + quantization.