← Back to leaderboard
local
/home/thomaskyn/models/Qwen3.6-35B-A3B/Qwen3.6-35B-A3B-UD-IQ3_S.gguf
Released Context 0Khome-thomaskyn-models-qwen3.6-35b-a3b-qwen3.6-35b-a3b-ud-iq3_s.gguf
PipelineScore
81.0MAINLINERanked #21 of 53 models · 60th percentileCode is the headline (100.0); throughput is the soft spot (17.3). Best-fit profile: Coding.
Category breakdown
Score per category, normalized 0–100 against the v1 anchor.
Code
100.0
Reason
80.0
Tool Use
87.5
RAG
100.0
Speed
17.3
Strengths
Code100.0
RAG100.0
Tool Use87.5
Sample tasks
A taste of what the test pack measures. Full prompts are private and rotated daily.
CodeDifficulty 1code-fib-1
Fibonacci function
Write a Python `fib(n)` returning the nth Fibonacci number, O(n).
ReasonDifficulty 1reason-math-1
Train meeting time
Two trains, opposite directions, given speeds and start times — when do they meet?
RAGDifficulty 2rag-extract-1
Extract metrics to JSON
From the context, extract net sales, operating margin, and free cash flow as a JSON object. Numbers only.
Tool UseDifficulty 2tool-schema-1
OpenAPI param selection
Given an OpenAPI schema with limit/offset/sort, fill JSON for 'next 50, recent first.'
RAGDifficulty 2rag-grounding-1
Refuses to fabricate
Context lacks the answer — does the model fabricate or correctly say it can't?
Compare with
/home/thomaskyn/models/Qwen3.6-35B-A3B/Qwen3.6-35B-A3B-UD-IQ3_S.gguf vs openai/gpt-oss-20b/home/thomaskyn/models/Qwen3.6-35B-A3B/Qwen3.6-35B-A3B-UD-IQ3_S.gguf vs DeepSeek R1 671B-A37B/home/thomaskyn/models/Qwen3.6-35B-A3B/Qwen3.6-35B-A3B-UD-IQ3_S.gguf vs Qwen 3 235B-A22B MoE/home/thomaskyn/models/Qwen3.6-35B-A3B/Qwen3.6-35B-A3B-UD-IQ3_S.gguf vs DeepSeek V3 671B-A37B/home/thomaskyn/models/Qwen3.6-35B-A3B/Qwen3.6-35B-A3B-UD-IQ3_S.gguf vs qwen3.6-35b-a3b