Back to search results

Qwen3-32B · Q4_K_M

NVIDIA RTX 309024GBllama.cpp b4567L1 Reproduced

This page aggregates 2 real-world runs of Qwen3-32B (Q4_K_M) on NVIDIA RTX 3090 with llama.cpp, contributed by 2 independent source platforms; metrics are averages of published measurements.

12.95 tok/s

Decode

Decode speed

143.35 tok/s

Prefill

Prefill speed

1.6 s

TTFT

Time to first token

19.5 GB

VRAM

VRAM usage

L1 Reproduced

2

Measured runs

2

Independent sources

GitHub, Bilibili

Source platforms

2026-08-15

Last verified

Performance

  1. llama.cpp · Q4_K_M (current)Decode 12.95 · Prefill 143.35 ·

Core figures

Decode (avg)
12.95 tok/s
Prefill (avg)
143.35 tok/s
TTFT (avg)
1.6 s
VRAM (avg)
19.5 GB
MTP acceptance rate
TTFB
— GB
Power draw
299.8 W

Configuration

Member-level fields are taken from the most recent run

Model
Qwen3-32B
Quantization
Q4_K_M
Framework
llama.cpp
Version
b4567
Context length
8192 tokens
GPU layers
99

Hardware

Nominal and measured figures are shown side by side; whether it runs is the reader's call

GPU
NVIDIA RTX 3090
Nominal VRAM
24 GB
Measured VRAM (avg)
19.5 GB
OS
Ubuntu 22.04
Driver
535.183.01
CUDA
12.2
Power draw
298.4 W

Sources & evidence

2 measured records in total, each traceable to its original source

  1. L0 Self-reportedBilibiliOriginal link Verified on 2026-07-30

    b4567 · Ubuntu 22.04 · CUDA 12.2 · 8192 ctx

    12.8 tok/s

    Decode

    141.7 tok/s

    Prefill

    1.62 s

    TTFT

    19.5 GB

    VRAM

    MTP

    298.4 W

    Power draw

    3090 上 Q4_K_M 约 13 tok/s。
  2. L1 ReproducedGitHubOriginal link Verified on 2026-08-15

    b4602 · Ubuntu 22.04 · CUDA 12.2 · 8192 ctx

    13.1 tok/s

    Decode

    145 tok/s

    Prefill

    1.58 s

    TTFT

    19.5 GB

    VRAM

    MTP

    301.2 W

    Power draw

    3090 复现 13.1 tok/s。