qwen3.5:35b-a3b-q4_K_M

qwen35moe · 36.0B · Q4_K_M

THINKING MODEL

Dell Inc. Precision Tower 7910 (Intel Xeon® E5-2637 v3)

96 GB · Microsoft Windows 10 Pro 10.0.19045

Tested on July 30, 2026 · Submitted by Csaba
Top 87% Compare
Global Score
49 /100
Not Rec.
Hardware Fit
54/100
Quality
47/100

Get this model

Hardware

Machine
Dell Inc. Precision Tower 7910
CPU
Intel Xeon® E5-2637 v3
Cores
16 total (16 perf)
Frequency
3.5 GHz
RAM
96 GB DDR4
GPU
NVIDIA GeForce RTX 5060 Ti
OS
Microsoft Windows 10 Pro 10.0.19045
Arch
x64
Power Mode
balanced

Performance

Tokens/sec
21.6
Standard deviation
±3.3
First chunk latency
1.5 s
Time to first token
30.0 s
Load time
53.8 s
Memory usage
22.7 GB (24%)
Total tokens
1429
Thinking tokens (est.)
~780

Score breakdown

Speed
24/50
Time to first token
0/20
Memory
30/30

Quality

Reasoning
14/20
Coding
1/20
Instruction following
6/20
Structured output
5/15
Math
11/15
Multilingual
10/10

Category levels

Reasoning: Adequate Coding: Poor Instruction Following: Weak Structured Output: Weak Math: Adequate Multilingual: Strong

Metadata

Spec version
0.2.1
Runtime
Ollama 0.32.1
Model format
GGUF
Hardware profile
HIGH-END
Result hash
dec36352b5527ca11f5870f86a18ec3ff67b76003c47afb9dde3c38640f61842

Interpretation

Hardware fit: 54/100. Overall suitability: NOT RECOMMENDED (Global 49/100). Category profile: Reasoning: Adequate, Coding: Poor, Instruction Following: Weak, Structured Output: Weak, Math: Adequate, Multilingual: Strong.

Disqualifiers

  • Time to first token too high: 30000ms (maximum: 13009ms for HIGH-END profile)

Bench Environment

CPU load: avg 31% (peak 40%)

Run yours now

$ npm install -g metrillm@latest
$ metrillm

Requires Node 20+ and Ollama or LM Studio running

Or run without installing: npx metrillm@latest