Benchmark LLMs on YOUR hardware.
The same 34 deterministic tasks, scored entirely on your machine — no judge model, no API key, one 0–100 score. A catalog of 109 models, ranked by runs people actually made on their own hardware. The only public LLM board that ranks where the model runs, not just which model it is.
Every score here is a real run. We seed nothing — which is why the board is small and why you can trust the numbers on it.
$ npx @pipelinescore/cliThat's the whole command. It finds your Ollama / LM Studio / llama.cpp / MLX server, lists your models, auto-detects your hardware, and walks you through the rest.
The model board
Best run per model. Click a column to re-rank, pick two rows to go head-to-head.
| # | vs | Model | PipelineScore ▾ | Tier |
|---|---|---|---|---|
| 1 | openai/gpt-oss-20blocal | 93.2 | TRUNK | |
| 2 | gpt-oss:latestlocal | 91.6 | TRUNK | |
| 3 | qwen3.6-35b-a3blocal | 87.5 | MAINLINE | |
| 4 | qwen3-coder:30blocal | 86.6 | MAINLINE | |
| 5 | qwen3.8:27blocal | 86.2 | MAINLINE | |
| 6 | qwen3.8-27b@iq3_xxslocal | 84.2 | MAINLINE | |
| 7 | mlx-community/Qwen3.6-35B-A3B-4bitlocal | 83.7 | MAINLINE | |
| 8 | /home/thomaskyn/models/Qwen3.6-35B-A3B/Qwen3.6-35B-A3B-UD-IQ3_S.gguflocal | 80.9 | MAINLINE | |
| 9 | gemma4:12b-it-qat_gpulocal | 80.9 | MAINLINE | |
| 10 | qwen/qwen3.6-35b-a3blocal | 79.3 | MAINLINE | |
| 11 | qwen2.5-coder:7blocal | 77.7 | MAINLINE | |
| 12 | /home/thomaskyn/models/Qwen3.5-4b/Qwen3.5-4B-Q4_K_M.gguflocal | 57.8 | TAP | |
| 13 | /home/thomaskyn/models/Qwen3.5-9b/Qwen3.5-9B-Q4_K_M.gguflocal | 50.5 | TAP |
Popular matchups
The rivalries worth settling. Every pair opens a live head-to-head.
Five measures. One number.
Code is executed, reasoning is exact-match, tool use and RAG are JSON-match, speed is measured throughput. No judge model, no rubric, no API key.
The tiers
Every score maps to a tier, named the way pipelines are: from TRUNK (top of the network) down to DRIP.
Point the CLI at your model
Ollama, LM Studio, MLX, llama.cpp — anything OpenAI-compatible. Local runs need no account and no API key.
Tag your hardware
--hardware-tag m3-max-128gb / rtx-4090-24gb / a100-80gb. Same model on different rigs gets ranked separately — that's the point.
Land on the board
A deterministic 0–100 PipelineScore across 34 tasks, computed on your machine, plus a tier badge and a public spot on the hardware-aware leaderboard.