Skip to content

Rankings

Who delivers, task by task.

Every workspace that reports outcomes (POST /api/v1/feedback) teaches this board. Targets are ranked per task domain and kind by the lower confidence bound of their reported success rate, so a long record beats a lucky streak; quality, latency and cost ride along. Aggregated across the deployment, never per workspace.

Targets ranked by reported outcomes over 90 days
#TargetScoreSuccessQualityRequestsp50 latencyCost / request
-
On-prem private modelllmprovisional
llm-onpremollama/qwen3:4b
0.342100.0%of 20.9019458%2.23 s$0.0018
-
Mid-tier modelllmprovisional
llm-midazure/llm-mid
0.206100.0%of 10.9051%35.0 ms$0
-
Small writing modelllmprovisional
writer-smallollama/gemma3:4b
--no outcomes yet-6118%1.0 min$0.0001
-
vision-ocrollama/qwen3-vl:4b
--no outcomes yet-237%2.4 min$0.0003
-
Translation skillskillprovisional
skill-translate
--no outcomes yet-175%25.0 ms$0
-
Small fast modelllmprovisional
llm-smallazure/llm-small
--no outcomes yet-165%45.0 ms$0.0000
-
Code modelllmprovisional
coder-smallollama/qwen2.5-coder:7b
--no outcomes yet-93%1.1 min$0.0001
-
llm-frontierazure/llm-frontier
--no outcomes yet-51%55.0 ms$0
-
SQL skillskillprovisional
skill-sql
--no outcomes yet-21%25.0 ms$0
-
gpt-ossollama/gpt-oss:20b
--no outcomes yet-10%35.0 ms$0
-
Human escalation queuehumanprovisional
human-escalationhuman/queue
--no outcomes yet-10%25.0 ms$0
-
Customer support agentagentprovisional
support-agentollama/qwen3:4b
--no outcomes yet-10%1.3 min$0.0007

335 routed requests over 90 days across every workspace · 0 of 12 targets ranked; the rest have fewer than 20 reported outcomes and are provisional. Score = Wilson lower bound (95%) of the reported success rate; p50 latency is read off a 10 ms histogram. No workspace, request or prompt data is exposed.

Raw numbers: /api/v1/rankings/targets and the task domains at /api/v1/rankings/targets/domains. Cross-check the reported outcomes against the traffic rankings and the published benchmarks.