Back to search results

Qwen3-32B · Q4_K_M

NVIDIA RTX 409024GBOllama 0.5.7L1 Reproduced

This page aggregates 1 real-world runs of Qwen3-32B (Q4_K_M) on NVIDIA RTX 4090 with Ollama, contributed by 1 independent source platforms; metrics are averages of published measurements.

17.4 tok/s

Decode

Decode speed

198.2 tok/s

Prefill

Prefill speed

1.31 s

TTFT

Time to first token

19.8 GB

VRAM

VRAM usage

L1 Reproduced

1

Measured runs

1

Independent sources

V2EX

Source platforms

23 days ago

Last verified

Performance

  1. llama.cpp · Q4_K_M Decode 18.4 · Prefill 216.93 ·
  2. Ollama · Q4_K_M (current)Decode 17.4 · Prefill 198.2 ·

Core figures

Decode (avg)
17.4 tok/s
Prefill (avg)
198.2 tok/s
TTFT (avg)
1.31 s
VRAM (avg)
19.8 GB
MTP acceptance rate
TTFB
— GB
Power draw
339 W

Configuration

Member-level fields are taken from the most recent run

Model
Qwen3-32B
Quantization
Q4_K_M
Framework
Ollama
Version
0.5.7
Context length
8192 tokens

Hardware

Nominal and measured figures are shown side by side; whether it runs is the reader's call

GPU
NVIDIA RTX 4090
Nominal VRAM
24 GB
Measured VRAM (avg)
19.8 GB
OS
Ubuntu 24.04
Driver
550.107.02
CUDA
12.4
Power draw
339 W

Sources & evidence

1 measured records in total, each traceable to its original source

  1. L1 ReproducedV2EXOriginal link Verified on 2026-08-28

    0.5.7 · Ubuntu 24.04 · CUDA 12.4 · 8192 ctx

    17.4 tok/s

    Decode

    198.2 tok/s

    Prefill

    1.31 s

    TTFT

    19.8 GB

    VRAM

    MTP

    339 W

    Power draw

    Ollama 默认参数下比 llama.cpp 直跑略低。