PipelineScore
LLM benchmarks · v3 testpack · deterministic · no API key

Benchmark LLMs on YOUR hardware.

The same 34 deterministic tasks, scored entirely on your machine — no judge model, no API key, one 0–100 score. A catalog of 109 models, ranked by runs people actually made on their own hardware. The only public LLM board that ranks where the model runs, not just which model it is.

Every score here is a real run. We seed nothing — which is why the board is small and why you can trust the numbers on it.

$ npx @pipelinescore/cli

That's the whole command. It finds your Ollama / LM Studio / llama.cpp / MLX server, lists your models, auto-detects your hardware, and walks you through the rest.

Models tracked 109Real runs 24Users 5Testpack v3Tasks 34

The model board

Best run per model. Click a column to re-rank, pick two rows to go head-to-head.

How scores are computed →
Weighting
13 models · sorted by Balanced composite (high → low) · pick any two rows to compare
#vsModelPipelineScore Tier
1openai/gpt-oss-20blocal
93.2
TRUNK
2gpt-oss:latestlocal
91.6
TRUNK
3qwen3.6-35b-a3blocal
87.5
MAINLINE
4qwen3-coder:30blocal
86.6
MAINLINE
5qwen3.8:27blocal
86.2
MAINLINE
6qwen3.8-27b@iq3_xxslocal
84.2
MAINLINE
7mlx-community/Qwen3.6-35B-A3B-4bitlocal
83.7
MAINLINE
8/home/thomaskyn/models/Qwen3.6-35B-A3B/Qwen3.6-35B-A3B-UD-IQ3_S.gguflocal
80.9
MAINLINE
9gemma4:12b-it-qat_gpulocal
80.9
MAINLINE
10qwen/qwen3.6-35b-a3blocal
79.3
MAINLINE
11qwen2.5-coder:7blocal
77.7
MAINLINE
12/home/thomaskyn/models/Qwen3.5-4b/Qwen3.5-4B-Q4_K_M.gguflocal
57.8
TAP
13/home/thomaskyn/models/Qwen3.5-9b/Qwen3.5-9B-Q4_K_M.gguflocal
50.5
TAP
vs

Popular matchups

The rivalries worth settling. Every pair opens a live head-to-head.

Five measures. One number.

Code is executed, reasoning is exact-match, tool use and RAG are JSON-match, speed is measured throughput. No judge model, no rubric, no API key.

Code
28%
Reason
22%
Tool Use
18%
RAG
17%
Speed
15%

The tiers

Every score maps to a tier, named the way pipelines are: from TRUNK (top of the network) down to DRIP.

TRUNK
90100
MAINLINE
7589
FEEDER
6074
TAP
4059
DRIP
039
STEP 01

Point the CLI at your model

Ollama, LM Studio, MLX, llama.cpp — anything OpenAI-compatible. Local runs need no account and no API key.

STEP 02

Tag your hardware

--hardware-tag m3-max-128gb / rtx-4090-24gb / a100-80gb. Same model on different rigs gets ranked separately — that's the point.

STEP 03

Land on the board

A deterministic 0–100 PipelineScore across 34 tasks, computed on your machine, plus a tier badge and a public spot on the hardware-aware leaderboard.