Skip to content

Rankings

Who delivers, task by task.

Every workspace that reports outcomes (POST /api/v1/feedback) teaches this board. Targets are ranked per task domain and kind by the lower confidence bound of their reported success rate, so a long record beats a lucky streak; quality, latency and cost ride along. Aggregated across the deployment, never per workspace.

Targets ranked by reported outcomes in General over 90 days
#TargetScoreSuccessQualityRequestsp50 latencyCost / request
-
On-prem private modelllmprovisional
llm-onpremollama/qwen3:4b
0.342100.0%of 20.901558%45.0 ms$0.0001
-
Mid-tier modelllmprovisional
llm-midazure/llm-mid
0.206100.0%of 10.9028%25.0 ms$0
-
Small fast modelllmprovisional
llm-smallazure/llm-small
--no outcomes yet-519%45.0 ms$0.0000
-
gpt-ossollama/gpt-oss:20b
--no outcomes yet-14%35.0 ms$0
-
Human escalation queuehumanprovisional
human-escalationhuman/queue
--no outcomes yet-14%25.0 ms$0
-
vision-ocrollama/qwen3-vl:4b
--no outcomes yet-14%3.0 min$0.0001
-
Small writing modelllmprovisional
writer-smallollama/gemma3:4b
--no outcomes yet-14%325.0 ms$0

26 routed requests over 90 days across every workspace · 0 of 7 targets ranked; the rest have fewer than 20 reported outcomes and are provisional. Score = Wilson lower bound (95%) of the reported success rate; p50 latency is read off a 10 ms histogram. No workspace, request or prompt data is exposed.

Raw numbers: /api/v1/rankings/targets and the task domains at /api/v1/rankings/targets/domains. Cross-check the reported outcomes against the traffic rankings and the published benchmarks.