PipelineScore
← Back to leaderboard
local

/home/thomaskyn/models/Qwen3.5-9b/Qwen3.5-9B-Q4_K_M.gguf

Released Context 0Khome-thomaskyn-models-qwen3.5-9b-qwen3.5-9b-q4_k_m.gguf
PipelineScore
50.5TAP
Tool use is the headline (87.5); throughput is the soft spot (12.0). Best-fit profile: Agentic.

Category breakdown

Score per category, normalized 0–100 against the v1 anchor.

Code
25.0
Reason
60.0
Tool Use
87.5
RAG
75.0
Speed
12.0

Strengths

Tool Use87.5
RAG75.0
Reason60.0

Sample tasks

A taste of what the test pack measures. Full prompts are private and rotated daily.

CodeDifficulty 1code-fib-1

Fibonacci function

Write a Python `fib(n)` returning the nth Fibonacci number, O(n).

ReasonDifficulty 1reason-math-1

Train meeting time

Two trains, opposite directions, given speeds and start times — when do they meet?

RAGDifficulty 2rag-extract-1

Extract metrics to JSON

From the context, extract net sales, operating margin, and free cash flow as a JSON object. Numbers only.

Tool UseDifficulty 2tool-schema-1

OpenAPI param selection

Given an OpenAPI schema with limit/offset/sort, fill JSON for 'next 50, recent first.'

RAGDifficulty 2rag-grounding-1

Refuses to fabricate

Context lacks the answer — does the model fabricate or correctly say it can't?

Compare with