Back to search results

Qwen3.6-27B · Q5_K_S

NVIDIA RTX 3090 Ti24GBllama.cpp v0.3.0-e0663be2713cL0 Self-reported

This page aggregates 1 real-world runs of Qwen3.6-27B (Q5_K_S) on NVIDIA RTX 3090 Ti with llama.cpp, contributed by 1 independent source platforms; metrics are averages of published measurements.

50.84 tok/s

Decode

Decode speed

1,376.9 tok/s

Prefill

Prefill speed

0.09 s

TTFT

Time to first token

22.47 GB

VRAM

VRAM usage

L0 Self-reported

1

Measured runs

1

Independent sources

GitHub

Source platforms

Today

Last verified

Performance

  1. llama.cpp · Q5_K_S (current)Decode 50.84 · Prefill 1376.9 ·

Core figures

Decode (avg)
50.84 tok/s
Prefill (avg)
1,376.9 tok/s
TTFT (avg)
0.09 s
VRAM (avg)
22.47 GB
MTP acceptance rate
TTFB
— GB
Power draw
392.82 W

Configuration

Member-level fields are taken from the most recent run

Model
Qwen3.6-27B
Quantization
Q5_K_S
Framework
llama.cpp
Version
v0.3.0-e0663be2713c
Context length
102400 tokens

Hardware

Nominal and measured figures are shown side by side; whether it runs is the reader's call

GPU
NVIDIA RTX 3090 Ti
Nominal VRAM
24 GB
Measured VRAM (avg)
22.47 GB
OS
Ubuntu 24.04.4 LTS
Driver
580.159.03
CUDA
13.0
Power draw
392.82 W

Sources & evidence

1 measured records in total, each traceable to its original source

  1. L0 Self-reportedGitHubOriginal link Verified on 2026-09-20

    v0.3.0-e0663be2713c · Ubuntu 24.04.4 LTS · CUDA 13.0 · 102400 ctx

    50.84 tok/s

    Decode

    1,376.9 tok/s

    Prefill

    0.09 s

    TTFT

    22.47 GB

    VRAM

    MTP

    392.82 W

    Power draw

    run-1 wall= 20.00s ttft= 80ms toks= 998 wall_TPS= 49.90 decode_TPS= 50.10 run-2 wall= 20.44s ttft= 91ms toks=1000 wall_TPS= 48.93 decode_TPS= 49.15 run-3 wall= 18.88s ttft= 81ms toks=1000 wall_TPS= 52.97 decode_TPS= 53.20 run-4 wall= 20.32s ttft= 83ms toks=1000 wall_TPS= 49.21 decode_TPS= 49.41 run-5 wall= 19.20s ttft= 92ms toks=1000 wall_TPS= 52.07 decode_TPS= 52.32