What AI can your machine run?

Query the model, framework, and quantization combos your hardware can run — based on real benchmark runs. Beyond speed: how reliable each data point is, whether it's directly comparable, and where the evidence comes from.

Trending searches

11

Records

Real benchmark runs

4

Hardware

Mainstream GPUs / CPUs

3

Models

LLMs & specialized models

3

Frameworks

llama.cpp / vLLM & more

L0 / L1 / L2

Evidence levels

Read the data notes →

Latest runs

The latest model run results from real collected data.

View all →

Qwen3-32B

L0 Self-reported

Q4_K_M · llama.cpp b4567

  • NVIDIA RTX 4090
  • 24GB
  • Ubuntu 24.04
  • CUDA 12.4

18tok/s

Decode

210.5tok/s

Prefill

19.5GB

VRAM

  • NVIDIA
  • llama.cpp
  • Self-reported
Source: GitHub2026-09-18

Qwen3-32B

L1 Reproduced

Q4_K_M · llama.cpp b4602

  • NVIDIA RTX 4090
  • 24GB
  • Ubuntu 24.04
  • CUDA 12.4

18.6tok/s

Decode

218tok/s

Prefill

19.4GB

VRAM

  • NVIDIA
  • llama.cpp
  • Community-verified
Source: GitHub2026-09-18

Qwen3-32B

L2 Cross-framework verified

Q4_K_M · llama.cpp b4602

  • NVIDIA RTX 4090
  • 24GB
  • Windows 11 24H2
  • CUDA 12.4

18.6tok/s

Decode

222.3tok/s

Prefill

19.6GB

VRAM

  • NVIDIA
  • llama.cpp
  • Cross-framework
Source: Reddit2026-09-18

Qwen3-32B

L1 Reproduced

Q4_K_M · Ollama 0.5.7

  • NVIDIA RTX 4090
  • 24GB
  • Ubuntu 24.04
  • CUDA 12.4

17.4tok/s

Decode

198.2tok/s

Prefill

19.8GB

VRAM

  • NVIDIA
  • Ollama
  • Community-verified
Source: V2EX2026-09-18

Why trust Arao AI?

A structured evidence system tells you how much each data point can be trusted.

More about the evidence system →
L0 Self-reported

Published by the community or individuals, not independently verified.

L1 Reproduced

Community reproductions — small-scale independent verification with reproduction counts.

L2 Cross-framework verified

Multi-framework benchmark runs on the same hardware + model + quantization.