Rankings
Which model, agent, skill, tool or person actually gets the job done, per kind of task - ranked by the outcomes people report (a thumbs up or down in the apps, POST /api/v1/feedback from code). The score is the lower confidence bound of the success rate, so a long record beats a lucky streak; quality, speed and cost ride along. Counted across the whole deployment, never shown per workspace. A row marked provisional has too few reports to trust yet.
| # | Target | Score | Success | Quality | Requests | p50 latency | Cost / request |
|---|---|---|---|---|---|---|---|
| - | embed-smallollama-qwen3-embedding-4b/qwen3-embedding:4b | - | -no outcomes yet | - |
295 routed requests over 30 days across every workspace · 0 of 9 targets ranked; the rest have fewer than 20 reported outcomes and are provisional. Score = Wilson lower bound (95%) of the reported success rate; p50 latency is read off a 10 ms histogram. No workspace, request or prompt data is exposed.
Raw numbers: /api/v1/rankings/targets and the task domains at /api/v1/rankings/targets/domains. Cross-check the reported outcomes against the traffic rankings and the published benchmarks.
| 11.2 s |
| $0 |
| - | writer-smallollama-qwen3-5-4b/qwen3.5:4b | - | -no outcomes yet | - | 5318% | 1.8 min | $0.0002 |
| - | llm-smallazure/llm-small | - | -no outcomes yet | - | 3913% | 1.92 s | $0.0005 |
| - | skill-translateollama-translategemma-4b/translategemma:4b | - | -no outcomes yet | - | 134% | 33.9 s | $0.0002 |
| - | coder-smallollama-glm-4-7-flash/glm-4.7-flash | - | -no outcomes yet | - | 62% | 2.7 min | $0.0004 |
| - | vision-ocrollama-glm-ocr/glm-ocr | - | -no outcomes yet | - | 62% | 14 min | $0.0006 |
| - | llm-midazure/llm-mid | - | -no outcomes yet | - | 31% | 2.10 s | $0.0021 |
| - | llm-onpremollama-qwen3-5-4b/qwen3.5:4b | - | -no outcomes yet | - | 31% | 2.6 min | $0.0048 |
| - | llm-frontierazure/llm-frontier | - | -no outcomes yet | - | 10% | 25.0 ms | $0 |