Skip to content

Rankings

Who delivers, task by task.

Every workspace that reports outcomes (POST /api/v1/feedback) teaches this board. Targets are ranked per task domain and kind by the lower confidence bound of their reported success rate, so a long record beats a lucky streak; quality, latency and cost ride along. Aggregated across the deployment, never per workspace.

Targets ranked by reported outcomes in Coding over 30 days
#TargetScoreSuccessQualityRequestsp50 latencyCost / request
-
On-prem private modelllmprovisional
llm-onpremollama/qwen3:4b
--no outcomes yet-3160%2.44 s$0.0023
-
Code modelllmprovisional
coder-smallollama/qwen2.5-coder:7b
--no outcomes yet-815%54.4 s$0.0002
-
Small writing modelllmprovisional
writer-smallollama/gemma3:4b
--no outcomes yet-713%1.0 min$0.0001
-
llm-frontierazure/llm-frontier
--no outcomes yet-48%55.0 ms$0
-
Small fast modelllmprovisional
llm-smallazure/llm-small
--no outcomes yet-12%585.0 ms$0.0000
-
SQL skillskillprovisional
skill-sql
--no outcomes yet-12%105.0 ms$0

52 routed requests over 30 days across every workspace · 0 of 6 targets ranked; the rest have fewer than 20 reported outcomes and are provisional. Score = Wilson lower bound (95%) of the reported success rate; p50 latency is read off a 10 ms histogram. No workspace, request or prompt data is exposed.

Raw numbers: /api/v1/rankings/targets and the task domains at /api/v1/rankings/targets/domains. Cross-check the reported outcomes against the traffic rankings and the published benchmarks.