Skip to content

Rankings

Who delivers, task by task.

Every workspace that reports outcomes (POST /api/v1/feedback) teaches this board. Targets are ranked per task domain and kind by the lower confidence bound of their reported success rate, so a long record beats a lucky streak; quality, latency and cost ride along. Aggregated across the deployment, never per workspace.

Targets ranked by reported outcomes in Math over 90 days
#TargetScoreSuccessQualityRequestsp50 latencyCost / request
-
On-prem private modelllmprovisional
llm-onpremollama/qwen3:4b
--no outcomes yet-10961%2.13 s$0.0018
-
Small writing modelllmprovisional
writer-smallollama/gemma3:4b
--no outcomes yet-4928%1.0 min$0.0001
-
vision-ocrollama/qwen3-vl:4b
--no outcomes yet-2011%2.3 min$0.0003

178 routed requests over 90 days across every workspace · 0 of 3 targets ranked; the rest have fewer than 20 reported outcomes and are provisional. Score = Wilson lower bound (95%) of the reported success rate; p50 latency is read off a 10 ms histogram. No workspace, request or prompt data is exposed.

Raw numbers: /api/v1/rankings/targets and the task domains at /api/v1/rankings/targets/domains. Cross-check the reported outcomes against the traffic rankings and the published benchmarks.