llama3.3:70b-instruct-q4_K_M

llama · 70.6B · Q4_K_M

Dell Inc. Precision Tower 7910 (Intel Xeon® E5-2637 v3)

96 GB · Microsoft Windows 10 Pro 10.0.19045

Tested on August 20, 2026 · Submitted by Csaba
Top 88% Compare
Global Score
51 /100
Not Rec.
Hardware Fit
31/100
Quality
60/100

Get this model

Hardware

Machine
Dell Inc. Precision Tower 7910
CPU
Intel Xeon® E5-2637 v3
Cores
16 total (16 perf)
Frequency
3.5 GHz
RAM
96 GB DDR4
GPU
NVIDIA GeForce RTX 3090, NVIDIA GeForce RTX 3060
OS
Microsoft Windows 10 Pro 10.0.19045
Arch
x64
Power Mode
balanced

Performance

Tokens/sec
1.7
Standard deviation
±0.0
First chunk latency
5.6 s
Time to first token
5.6 s
Load time
0.4 s
Memory usage
50.3 GB (53%)
Total tokens
389

Score breakdown

Speed
2/50
Time to first token
5/20
Memory
24/30

Quality

Reasoning
17/20
Coding
6/20
Instruction following
6/20
Structured output
11/15
Math
11/15
Multilingual
9/10

Category levels

Reasoning: Strong Coding: Weak Instruction Following: Weak Structured Output: Adequate Math: Adequate Multilingual: Strong

Metadata

Spec version
0.2.1
Runtime
Ollama 0.32.1
Model format
GGUF
Hardware profile
HIGH-END
Result hash
fc925a43e4145dde354f9f3f1df24a95d7cd86ab845df109221f8ad0d7353e10

Interpretation

Hardware fit: 31/100. Overall suitability: NOT RECOMMENDED (Global 51/100). Category profile: Reasoning: Strong, Coding: Weak, Instruction Following: Weak, Structured Output: Adequate, Math: Adequate, Multilingual: Strong.

Disqualifiers

  • Token speed too low: 1.7 tok/s (minimum: 7 tok/s for HIGH-END profile)

Bench Environment

CPU load: avg 35% (peak 42%)

Run yours now

$ npm install -g metrillm@latest
$ metrillm

Requires Node 20+ and Ollama or LM Studio running

Or run without installing: npx metrillm@latest