Back to search results

Qwen3.8-27B · Q4_K_M

AMD Radeon RX 7900 GRE + NVIDIA GTX 1080Ti27GBllama.cpp 10453L0 Self-reported

This page aggregates 1 real-world runs of Qwen3.8-27B (Q4_K_M) on AMD Radeon RX 7900 GRE + NVIDIA GTX 1080Ti with llama.cpp, contributed by 1 independent source platforms; metrics are averages of published measurements.

11.13 tok/s

Decode

Decode speed

71.46 tok/s

Prefill

Prefill speed

s

TTFT

Time to first token

17.2 GB

VRAM

VRAM usage

L0 Self-reported

1

Measured runs

1

Independent sources

Reddit

Source platforms

Today

Last verified

Performance

  1. llama.cpp · Q6_K Decode 12.53 · Prefill 116.21 ·
  2. llama.cpp · Q4_K_M (current)Decode 11.13 · Prefill 71.46 ·

Core figures

Decode (avg)
11.13 tok/s
Prefill (avg)
71.46 tok/s
TTFT (avg)
— s
VRAM (avg)
17.2 GB
MTP acceptance rate
TTFB
— GB
Power draw
— W

Configuration

Member-level fields are taken from the most recent run

Model
Qwen3.8-27B
Quantization
Q4_K_M
Framework
llama.cpp
Version
10453
Flash Attention
On

Hardware

Nominal and measured figures are shown side by side; whether it runs is the reader's call

GPU
AMD Radeon RX 7900 GRE + NVIDIA GTX 1080Ti
Nominal VRAM
27 GB
Measured VRAM (avg)
17.2 GB
OS
Ubuntu

Get started

Startup command

Command from the latest benchmark run
llama/llama-bench --rpc 10.0.0.75:50053 -m /Qwen3.8-27B-Q4_K_M.gguf -fa on

Sources & evidence

1 measured records in total, each traceable to its original source

  1. L0 Self-reportedRedditOriginal link Verified on 2026-09-20

    10453 · Ubuntu

    11.13 tok/s

    Decode

    71.46 tok/s

    Prefill

    s

    TTFT

    17.2 GB

    VRAM

    MTP

    W

    Power draw

    RPC Radeon plus GTX 1080Ti Q4\_K using RPC |model|size|params|test|t/s| |qwen35 27B Q4\_K - Medium|15.92 GiB|27.32 B|pp512|71.46 ± 0.11| |qwen35 27B Q4\_K - Medium|15.92 GiB|27.32 B|tg128|11.13 ± 3.99| build: 3cb7ffb1a (10453) about 17.2GB VRAM but includes desktop resources about 2gb combined