/home/thomaskyn/models/Qwen3.5-4b/Qwen3.5-4B-Q4_K_M.gguf
Category breakdown
Score per category, normalized 0–100 against the v1 anchor.
Strengths
Same model, different rigs
Every submission of /home/thomaskyn/models/Qwen3.5-4b/Qwen3.5-4B-Q4_K_M.gguf on the 0–100 scale. The spread is the point: where it runs changes what you get.
Best 57.8 on gtx-1060-5gb-5gb · lowest 57.8 on gtx-1060-5gb-5gb · spread 0.0 pts across 2 runs. Hover a dot for its rig.
Sample tasks
A taste of what the test pack measures. Full prompts are private and rotated daily.
Fibonacci function
Write a Python `fib(n)` returning the nth Fibonacci number, O(n).
Train meeting time
Two trains, opposite directions, given speeds and start times — when do they meet?
Extract metrics to JSON
From the context, extract net sales, operating margin, and free cash flow as a JSON object. Numbers only.
OpenAPI param selection
Given an OpenAPI schema with limit/offset/sort, fill JSON for 'next 50, recent first.'
Refuses to fabricate
Context lacks the answer — does the model fabricate or correctly say it can't?