<!-- OpenSmartRoute: Benchmarks and leaderboard recipe. Source https://github.com/isathish/OpenSmartRoute/blob/main/examples/leaderboard/README.md; HTML https://opensmartroute.ai/docs/leaderboard -->
# Reproducible leaderboard runs

Everything the public-suite numbers in [docs/ROADMAP.md](https://opensmartroute.ai/docs/ROADMAP.md) (v0.4 measurement) are
made of, as commands you can re-run. The same recipe is what an OpenSmartRoute entry on RouterArena,
LLMRouterBench or xRouteBench is built from; the runs themselves are tracked under v1.0 in the roadmap and
recorded, with their submission status, in [results/README.md](https://opensmartroute.ai/docs/leaderboard-results.md#listings).

Requirements: `pip install "opensmartroute[yaml]"`, network access to the Hugging Face datasets-server
(public suites only; RouterBench / RouterEval are gated and need a Hub token plus `--preset`).

## 1. Catalogue

[targets.yaml](https://github.com/isathish/OpenSmartRoute/blob/main/examples/leaderboard/targets.yaml): three tiers, `llm-small` / `llm-mid` / `llm-frontier`. Battle datasets name
2023-24 models; `osr collect --tier` maps each model to the tier it played in (`eval.collect.ARENA_TIERS`),
so every row's label is "which tier was good enough for this prompt" - the RouteLLM framing.

## 2. Collect the public suites

```bash
osr collect --cache-dir data --limit 4000 \
    --tier frontier=llm-frontier --tier mid=llm-mid --tier small=llm-small
```

Writes one `<source>-tiers.jsonl` (+ `.meta.json`) per public source: `arena-55k`, `arena-100k`,
`arena-140k`, `routellm-battles`, `ppe-human`, `webdev-arena`, `mt-bench-human`, `reward-bench`,
`ultrafeedback`, `routellm-gpt4`. `--source NAME` restricts the run; `--refresh` ignores the cache.

## 3. Train the routing SLM on the training suites only

```bash
osr -t examples/leaderboard/targets.yaml slm train --out slm.json --cache-dir data \
    --source arena-55k-tiers --source routellm-battles-tiers --report --seed 0
```

`--report` holds out 10 % and prints its accuracy. The suites in step 4 are **not** in `--source`, so
they are held out at the suite level, not just the row level.

## 4. Evaluate on the held-out suites

```bash
for suite in mt-bench-human ppe-human webdev-arena; do
  osr -t examples/leaderboard/targets.yaml --slm slm.json eval data/$suite-tiers.jsonl \
      --baselines --calibration --robustness --limit 1000 --json > results/$suite.json
done
# declarative router (no learned signals) for the ablation column
osr -t examples/leaderboard/targets.yaml eval data/mt-bench-human-tiers.jsonl --baselines --limit 1000
```

What each flag reports and where it lands in the table:

| Flag | Reports | Table column |
|---|---|---|
| (default) | accuracy, mean quality, cost, latency, mean confidence, raw ECE / Brier | Router, ECE raw |
| `--baselines` | random / cheapest / best-prior / static task table / oracle, gap to oracle, noise floor | Static task table, Random, Cheapest |
| `--robustness` | repeat flip rate, paraphrase stability, profile-swap dependence (first 200 rows) | Paraphrase |
| `--calibration` | ECE, Brier, reliability bins, and `held_out`: temperature and isotonic calibrators fitted on the even rows, scored on the odd rows | ECE held-out |

The published table used `--limit 1000` per suite and seed 0 throughout; the SLM bundle and the
`Router` are deterministic, so re-running gives the same numbers for the same cache.

## 5. Gated suites (RouterBench, RouterEval, xRouteBench, RouterXBench)

Download the benchmark file with your Hub token, then evaluate it directly through the layout preset:

```bash
osr -t examples/leaderboard/targets.yaml --slm slm.json eval routerbench_0shot.pkl.jsonl \
    --preset routerbench --baselines --calibration --limit 5000
```

These rows carry several sampled answers per model, so `--baselines` can also report the label-noise
floor (`noise_floor()`), which the pairwise suites above cannot.

## 6. Publish

Commit `results/*.json`, the exact `osr --version`, the dataset ids, splits and row counts, and the
SHA-256 of `slm.json` together; that triple is the reproducible config a leaderboard entry links to.
`python scripts/exit_criteria.py --full` prints the synthetic exit criteria alongside, so the two
tables can be published from one run.

The run published with 0.5.0 is in [results/README.md](https://opensmartroute.ai/docs/leaderboard-results.md) (one JSON per suite, learned
and declarative); it is the entry submitted to RouterArena, LLMRouterBench and xRouteBench.
