hf.co/InternScience/Agents-A1-4B-Q8_0-GGUF:Q8_0

qwen35 · 4.21B · unknown

THINKING MODEL

MacBook Pro (Apple M1 Pro)

16 GB · macOS 26.6.2

Tested on August 22, 2026
Top 81% Compare
Global Score
59 /100
Marginal
Hardware Fit
97/100
Quality
43/100

Get this model

Hardware

Machine
MacBook Pro
CPU
Apple M1 Pro
Cores
10 total (8 perf + 2 eff)
Frequency
2.4 GHz
RAM
16 GB LPDDR5
GPU
Apple M1 Pro
OS
macOS 26.6.2
Arch
arm64
Power Mode
balanced

Performance

Tokens/sec
30.9
Standard deviation
±0.0
First chunk latency
891 ms
Time to first token
891 ms
Load time
3.6 s
Memory usage
6.6 GB (41%)
Total tokens
2634

Score breakdown

Speed
50/50
Time to first token
20/20
Memory
27/30

Quality

Reasoning
13/20
Coding
15/20
Instruction following
2/20
Structured output
0/15
Math
13/15
Multilingual
0/10

Category levels

Reasoning: Adequate Coding: Adequate Instruction Following: Poor Structured Output: Poor Math: Strong Multilingual: Poor

Metadata

Spec version
0.2.1
Runtime
Ollama 0.32.14
Model format
GGUF
Hardware profile
ENTRY
Result hash
cbda9e8bd44d9db3542abee159026c091a8e98d794d7dd5a8f76b929c3ccfba6

Interpretation

Hardware fit: 97/100. Overall suitability: MARGINAL (Global 59/100). Category profile: Reasoning: Adequate, Coding: Adequate, Instruction Following: Poor, Structured Output: Poor, Math: Strong, Multilingual: Poor.

Bench Environment

Power: AC CPU load: avg 32% (peak 61%)

Run yours now

$ npm install -g metrillm@latest
$ metrillm

Requires Node 20+ and Ollama or LM Studio running

Or run without installing: npx metrillm@latest