Back to search results
GLM-5.3-Flash · UD-IQ4_XS
2x NVIDIA RTX 509064GBllama.cpp moecachev1.6-rc0L0 Self-reported
This page aggregates 1 real-world runs of GLM-5.3-Flash (UD-IQ4_XS) on 2x NVIDIA RTX 5090 with llama.cpp, contributed by 1 independent source platforms; metrics are averages of published measurements.
21.96 tok/s
Decode
Decode speed
421.07 tok/s
Prefill
Prefill speed
0.6 s
TTFT
Time to first token
60.3 GB
VRAM
VRAM usage
L0 Self-reported
1
Measured runs
1
Independent sources
GitHub
Source platforms
Today
Last verified
Performance
- llama.cpp · UD-IQ4_XS (current)Decode 21.96 · Prefill 421.07 ·
Core figures
- Decode (avg)
- 21.96 tok/s
- Prefill (avg)
- 421.07 tok/s
- TTFT (avg)
- 0.6 s
- VRAM (avg)
- 60.3 GB
- MTP acceptance rate
- 0.64
- TTFB
- — GB
- Power draw
- — W
Configuration
Member-level fields are taken from the most recent run
- Model
- GLM-5.3-Flash
- Quantization
- UD-IQ4_XS
- Framework
- llama.cpp
- Version
- moecachev1.6-rc0
- Context length
- 204800 tokens
- Batch size
- 2048
- GPU layers
- 99
Hardware
Nominal and measured figures are shown side by side; whether it runs is the reader's call
- GPU
- 2x NVIDIA RTX 5090
- Nominal VRAM
- 64 GB
- Measured VRAM (avg)
- 60.3 GB
- OS
- Ubuntu 24.04.4 LTS
- Driver
- 610.57.04
- CUDA
- 13.3
Get started
Startup command
Command from the latest benchmark runmodel=/models/glm-5.3-flash-gguf/UD-IQ4_XS/GLM-5.3-Flash-UD-IQ4_XS-00001-of-00005.gguf offload=blk\.[0-9]+\.ffn_(gate|up|down)_exps\.weight=CPU ubatch=2048 ngl=99 split=1,1Sources & evidence
1 measured records in total, each traceable to its original source
moecachev1.6-rc0 · Ubuntu 24.04.4 LTS · CUDA 13.3 · 204800 ctx
21.96 tok/s
Decode
421.07 tok/s
Prefill
0.6 s
TTFT
60.3 GB
VRAM
0.64
MTP
— W
Power draw
[bench] narrative run 1/5: 1000 tok in 47.2s (21.42 decode TPS, ttft 514ms) [bench] narrative run 2/5: 1000 tok in 45.3s (22.41 decode TPS, ttft 630ms) [bench] narrative run 3/5: 1000 tok in 46.9s (21.55 decode TPS, ttft 538ms) [bench] narrative run 4/5: 1000 tok in 47.3s (21.49 decode TPS, ttft 747ms) [bench] narrative run 5/5: 1000 tok in 44.2s (22.93 decode TPS, ttft 550ms) [bench] prompt-processing run 1/1: 9989 prompt tok, 421.1 PP tok/s, ttft 23723ms