Roadmap
What is complete, what is open and the exit criteria for each milestone.
Status values: Shipped (in a tagged release), Complete (every item on main and the exit
criterion measured and met; ships with the next tag), In progress (code on main, part of the
exit criterion still open), Planned, Research (tracked, not scheduled). Where a criterion is
open, the section names exactly what is missing and what would close it.
Versioning follows SemVer. Each release has an exit criterion that must be met before tagging. Items are ordered by dependency, not by preference. Dates are deliberately absent; the project ships when the criteria are met.
This document is rebuilt from two inputs: (a) an engineering audit of the v0.3.0 code base
(what is implemented, what is heuristic, what is missing; section 0), and (b) a literature scan of LLM /
agent / tool routing up to September 2026 (summarised as design principles in section 0.3, cited per
item below, catalogued in RESEARCH.md). Every planned item names the paper or
concept it operationalises so the why survives the what.
v0.4.0 shipped first implementations across the v0.4 to v0.9 themes and v0.5.0 the research track,
the routing SLM and autopilot, estimates and the MCP server, the CLI sign-in flow and end-to-end tracing
(see CHANGELOG.md).
The research track that followed (verbalised difficulty, draft-response and hidden-state
representations, elastic budgets, personalisation, self-escalation, hand-off, aggregation, protocol
selection, schema-aware and cache-preserving discovery, skill graphs, MCP server recommendation,
semantic-router import, energy models, edge/cloud and auctions) is implemented on main with tests;
each row below names the module. The per-version sections keep their exit criteria; an item is only
Shipped when its code is released, and a version is only closed when its exit criterion has been
measured and published.
0. Where we were: engineering audit of v0.3.0#
0.1 What exists (and how it works)#
| Layer | Implemented | Mechanism |
|---|---|---|
| Signals | LengthSignal, LanguageSignal, ModalitySignal, ComplexitySignal, DomainActionSignal, PII / jailbreak regexes | Lexicon + regex heuristics, <1 ms, zero deps. Nothing is learned. |
| Policy | Policy.check/filter | Hard constraints before scoring: enable, allow/deny, kind, region, boundary, PII, tenant, cost/latency/token SLO, modality, language, jailbreak ≥ 0.5 → safety/human only |
| Strategies | Rules (pin/prefer/avoid), Capability, Similarity (hashing embedder, optional sentence-transformers), Bandit (Beta–Bernoulli TS per domain), LLMJudge, Cascade (executor, not a scorer) | Ensemble $\hat q=\sum_k w_k\kappa_k s_k / \sum_k w_k\kappa_k$; utility $U=w_q\hat q-w_c\tilde c-w_\ell\tilde\ell$ with log-min-max normalisation and a hard quality floor; confidence = softmax margin at fixed $\tau=0.1$ |
| Learning | IRTStrategy (2PL, online SGA), PreferenceStrategy (Bradley–Terry / Elo per domain), LinUCBStrategy (disjoint LinUCB on a 8+ | domains |
| Maths | ThompsonBeta (with decay), UCB1, LinUCB, EpsilonGreedy, CostAwareBandit (Lagrangian dual ascent), IRTModel, BradleyTerry, Elo, MarkovChain, RoutingMDP, EWMA, Welford, PageHinkley, WindowDrift, Wilson, ECE, Pareto, TOPSIS, Erlang‑C, Little, Kingman | Pure Python, deterministic, serialisable |
| Real-time | CircuitBreaker, TokenBucket, Budget, HealthRegistry/Policy/Strategy | Non-consuming checks in policy; one token reserved for the chosen target only |
| Enterprise | RouterBuilder, EnterpriseRouter, middleware (Guard, Tenant, Cache, Timeout), StateStore (memory, file), Telemetry (logging, metrics), AuditSink (hash-chained JSONL) | Ports are ABCs; only in-process / file adapters ship |
| Security | InputGuard (size, token estimate, head/tail naturalness gadget score), sanitize_for_prompt, Redactor, GuardMiddleware, load_secret | Threat model T1–T10 in SECURITY.md |
| Adapters | OpenAICompatClient (stdlib, retries, typed errors), judge_fn, embedder, chat_handler, agent harnesses (CallableHarness, HTTPHarness, SubprocessHarness → HarnessResult), SKILL.md loader (load_skills, agentskills.io frontmatter → RouteTarget), sentence-transformers, OpenTelemetry | No connection pooling |
| Execution | execution.execute/aexecute, Router.run, AsyncRouter.run, EnterpriseRouter.run | Plan-aware: skill pre-processor → persona/skill/primary instructions composed into context["system"] → primary handler; one Outcome per participant with role and shared task_id; slots gated by slot_quality_floor |
| Eval | evaluate, cost_quality_frontier, area_under_frontier, expected_calibration_error | Offline, single-draw labels, no counterfactual estimation |
| Surface | Router, AsyncRouter, FastAPI create_app, osr CLI, config.load_targets/load_rules, FeedbackStore (+ RouteLLM preference export) |
0.2 What was missing (from the audit) and where it landed#
| Gap | Where it bit | Roadmap home | Status |
|---|---|---|---|
| All signals are lexicon heuristics; no learned difficulty / task-type / reasoning-need classifier | signals/__init__.py | v0.4 | Shipped in 0.4.0 (signals.models, signals.ontology) |
No static task-type table baseline or "always-strongest" baseline in osr eval | eval/__init__.py | v0.4 | Shipped in 0.4.0 (eval.baselines, TaskTableStrategy) |
| Single-draw oracle in eval; no paraphrase-robustness or pool-diversity metrics | eval/__init__.py | v0.4 | Shipped in 0.4.0 (multi_sample_oracle, eval.robustness) |
| Confidence uses fixed $\tau$; ECE computed but never used to recalibrate; no abstention / set-valued output | router.py::_softmax_margin | v0.4, v0.8 | Shipped in 0.4.0 (math.calibration, RouteDecision.candidate_set) |
| Target = one model; reasoning effort / think-mode is not a routable dimension | core/types.py | v0.5 | Shipped in 0.4.0 (RouteTarget.effort, EffortStrategy) |
| No personalisation: user / tenant profile does not condition scores | core/types.py, strategies | v0.5 | Shipped in 0.4.0 (RouteRequest.profile, ProfileSignal) |
Cascade runs outside the ensemble and its steps are not fed back to learners; no "stop" action; fixed order | strategies/cascade.py | v0.6 | Shipped in 0.4.0 (CascadePlanner, BeliefTracker) |
| No trajectory step index; learners treat each outcome as immediate, so delayed, task-level reward is not joined | core/types.py::Outcome, learning/ | v0.6 | Shipped in 0.4.0 (Outcome.step, TaskCredit, TaskPins) |
| Plan slots emit outcomes but strategies do not condition on role | strategies/, learning/ | v0.6 | Shipped in 0.4.0 (signals.extra["role"], role-keyed learner memory: strategies.base.role_of / memory_key) |
No MCP / A2A import; no retrieve-then-rank narrowing for large catalogues; SimilarityStrategy is flat top-k | adapters/ | v0.7 | Shipped in 0.4.0 (adapters.mcp, adapters.a2a, retrieval) |
Only in-memory / file StateStore; per-outcome writes; no schema version in saved state | enterprise/__init__.py, learning/__init__.py | v0.7, v0.8 | Shipped in 0.4.0 (enterprise.stores) |
IRT / BT / LinUCB / Markov have no forgetting; only ThompsonBeta decays | math/irt.py, math/preference.py, math/bandits.py, math/markov.py | v0.8 | Shipped in 0.4.0 (decay= / discount= on every learner) |
CostAwareBandit handles one budget; no multi-resource knapsack, no shadow-audit stream | math/bandits.py | v0.8 | Shipped in 0.4.0 (MultiKnapsackBandit) |
Latency is a static latency_ms; queue state / TTFT is not modelled although Erlang-C exists | realtime/, math/decision.py | v0.8 | Shipped in 0.4.0 (enterprise.ops.QueueAwareStrategy) |
| No off-policy evaluation (IPS / DR), no shadow mode, no A/B harness | eval/ | v0.8 | Shipped in 0.4.0 (eval.ope, enterprise.ops.ShadowMiddleware) |
No learning-to-defer objective for TargetKind.HUMAN | strategies | v0.8 | Shipped in 0.4.0 (DeferStrategy) |
bandits._inverse is a naive Gauss-Jordan inverse (can go singular); posterior variance not surfaced in StrategyScore | math/bandits.py | v0.8 | Shipped in 0.4.0 (Sherman-Morrison updates, variance) |
| Energy / carbon not a cost dimension | core/types.py::cost | v0.8 | Shipped in 0.4.0 (unit_energy, unit_carbon, Objective(energy=, carbon=)) |
| Gadget detector is heuristic; no origin / provenance check on tool parameters in plans; no federated learning of router weights | security/ | v0.9 | Shipped in 0.4.0 (GadgetDetector, OriginPolicy, merge() on every learner + learning.merge_learners) |
| Untested: cascade path, audit-chain verification, judge JSON failure modes, state-corruption recovery, multi-replica consistency | tests/ | every release | Shipped in 0.4.0 (tests/test_resilience.py, test_cascade_*; FileAuditSink.verify, AutoLearner.quarantined / refresh) |
0.3 Evidence-driven principles (literature, 2025–2026)#
These findings change defaults, not just features. Each is cited in RESEARCH.md.
- Beat the boring baselines first. Under unified evaluation many routers, including commercial ones,
fail to reliably outperform simple baselines, and a large router-to-oracle gap remains (LLMRouterBench,
Jan 2026). A static task-type → model table captured 21 of 29 routable questions in one study
("Most of the LLM Routing Gap Is Task Type", Aug 2026). ⇒
osr evalmust always report always-strongest, always-cheapest and static task table alongside any learned strategy, and aTaskTableStrategybecomes a first-class, cheapest-possible learner. - The oracle is noisy. 12–36 % of the reported oracle gap on open-model pools is single-draw label noise that no single-commit router can recover ("How Much of the Routing Gap Is Real?", Jul 2026). ⇒ evaluation adopts a multi-sample oracle protocol and reports a reproducible headroom figure.
- Curate, don't pile. Larger ensembles show diminishing returns versus careful curation (LLMRouterBench); hierarchic social entropy saturates below ~10 behaviourally distinct actors ("When is Routing Meaningful?", Jul 2026). ⇒ ship a diversity / coreset report for catalogues.
- Robustness is a metric. Surface-form paraphrases should route identically; kNN-style routers
collapse under perturbation while prompted routers stay stable (same paper). ⇒ a paraphrase
robustness score in
osr eval. - Route the task, not the call. Agentic workflows have one delayed, task-level outcome; per-call
routers mis-attribute feedback (TRACE-Router, Jul 2026; MTRouter, ACL 2026; ProgRouter, EMNLP 2026).
⇒
Outcomegainstask_id; admission-time routing with pinned backend; terminal reward. - Escalation is an optimal-stopping problem, not a threshold (Bayesian self-escalation, Aug 2026; TACIT-Switch, Aug 2026; RLCascadeRouter, Aug 2026). ⇒ cascade becomes an MDP with a stop action.
- Control risk, not just cost. Conformal / distribution-free calibration gives set-valued routing
with abstention and provable mis-routing bounds (RACER, Feb 2026; RouteNLP, ACL 2026 industry;
CR², May 2026). ⇒ conformal thresholds replace hand-tuned
escalate_below. - Budgets are knapsacks and the world drifts. Multi-resource constraints, shadow-audit streams,
pessimistic reward / optimistic cost, hard meters before commitment (Drift-Aware Sparse Routing,
Sep 2026). ⇒ generalise
CostAwareBandit; add forgetting to every learner. - Latency is queue state, not a constant (Latency-aware routing with a TTFT estimator, Jul 2026;
HW-Router, DAC 2026). ⇒ wire
math.decisionqueueing intoHealthStrategy. - Discovery at scale is retrieve-then-rank. In-context selection collapses from Match@1 0.85 → 0.12 as registries grow to thousands of models / agents / tools / skills; retrieve-then-rank crosses over at N≈500 (Enrich-Retrieve-Rank, Aug 2026); hybrid BM25 + dense with RRF cut MCP tool tokens by 99 % in production (SCOUT, Aug 2026). ⇒ two-stage catalogue narrowing before scoring.
- Skill sets are submodular knapsacks (Best Prefix Selection, Aug 2026). ⇒
RoutePlanskill slots are selected as a set under a token budget, not top‑k. - Reasoning effort is a routable dimension (Think When Needed, SIGIR 2026; SCX Router predicts
task type, difficulty, reasoning mode and output length before generation, Sep 2026). ⇒ targets
carry
effortvariants; signals predict whether to think. - Personalisation matters exactly when it changes the answer (GMTRouter, EMNLP 2026; SkillFeed profile-conditioned skill routing, Aug 2026; xRouteBench personalised tasks). ⇒ profile-conditioned scoring with counterfactual tests.
- Routing data is fragmented and private. Federated router training improves the accuracy–cost frontier over client-local routers ("Federate the Router", Jan 2026). ⇒ mergeable learner state.
- The control plane is an attack surface. Beyond confounder gadgets: indirect prompt injection via tool content (ROPE, Aug 2026), unsigned MCP manifests, and recognition-without-enforcement gaps. ⇒ origin checks on state-changing tool parameters inside routed plans.
v0.3: Real integrations (Shipped)#
Goal: route to real providers with the same code that runs in tests.
| Item | Status | Notes |
|---|---|---|
| OpenAI-compatible HTTP client (stdlib, retries, typed errors) | Shipped | adapters.OpenAICompatClient; works with OpenAI, Azure, vLLM, Ollama, LiteLLM, OpenRouter, Groq |
| Judge / embedder / handler glue for the client | Shipped | judge_fn, embedder, chat_handler |
| sentence-transformers embedder | Shipped | adapters.sentence_transformers_embedder (extra: embeddings) |
| OpenTelemetry telemetry adapter | Shipped | adapters.OpenTelemetryTelemetry |
| Public config loaders with validation errors | Shipped | config.load_targets / load_rules |
| Rate-limit tokens consumed only for the chosen target | Shipped | bug fix in HealthPolicy / EnterpriseRouter |
Exit criterion: an end-to-end test routes, escalates to an LLM judge, and executes against an in-process OpenAI-compatible server with zero mocks of the transport. Met.
v0.4: Learned signals and honest evaluation (Shipped)#
Goal: replace lexicon heuristics with trained, still-tiny models, without adding torch to the core,
and make osr eval hard to fool.
Ordering rationale: principle 1 (baselines) and 2 (noisy oracle) come first because they decide whether anything in v0.5+ is worth building.
0.4a Evaluation harness#
| Item | Status | Notes |
|---|---|---|
Mandatory baselines in osr eval: always-strongest, always-cheapest, random, static task-type table, oracle | Shipped | osr eval --baselines; eval.baselines reports gap-to-oracle and the label-noise floor (LLMRouterBench 2601.07206; "Most of the routing gap is task type" 2608.23023) |
| Multi-sample oracle protocol | Shipped | multi_sample_oracle() and noise_floor() decompose the gap into label-noise floor vs recoverable (2607.03436) |
| Run-to-run variance | Shipped | repeat_flip_rate() in osr eval --robustness (5 % run-to-run flips observed in 2608.23023) |
| Paraphrase-robustness score | Shipped | paraphrase_robustness(): deterministic surface-form rewrites, % identical decisions (2607.09197) |
| Pool-diversity report | Shipped | eval.robustness.diversity() and greedy k-center coreset() |
| Latency-aware frontier | Shipped | eval.frontier.frontier3() and hypervolume(); osr eval --frontier3 |
| Dataset adapters | Shipped | eval.datasets presets: routerbench, llmrouterbench, routereval, xroutebench, routerxbench; osr eval --preset |
| Component-level ablation report | Shipped | ablation_report() drops one strategy or extractor at a time; osr eval --ablation |
0.4b Learned signals (zero core deps)#
| Item | Status | Notes |
|---|---|---|
TaskTableStrategy | Shipped | per-(task type x target) success table with Wilson intervals; fit_task_table() from outcomes |
| Task-type classifier | Shipped | TaskTypeSignal: multinomial logistic regression on hashed n-grams over signals.ontology.TaskOntology (SCX Router 2609.02292); weights ship as JSON; falls back to DOMAIN_LEXICON |
| Difficulty classifier | Shipped | LearnedDifficultySignal; training on RouterBench / LLMRouterBench outcomes via osr train --from |
| Reasoning-need signal | Shipped | ReasoningNeedSignal (Think When Needed 2601.18146); consumed by EffortStrategy |
| Expected-output-length signal | Shipped | OutputLengthSignal regression head |
| Verbalised-difficulty retrieval | Shipped | signals.VerbalisedDifficultySignal: reads a small model's verbalised difficulty (number, 7/10, or words) from context["difficulty"] or a callable, blends it into complexity and raises reasoning_need; parse_difficulty() (VDAR-Router 2607.18098) |
Training CLI osr train --from feedback.jsonl | Shipped | produces the SignalModelBundle JSON; --gadget also trains the gadget detector |
| Cold-start synthetic supervision | Shipped | synthesize_dataset() generates labelled prompts per ontology node (TRouter 2604.09377; SCX); osr train --synth |
0.4c Calibration#
| Item | Status | Notes |
|---|---|---|
Temperature / isotonic fitting of confidence | Shipped | TemperatureScaler, IsotonicCalibrator; osr eval --calibration reports ECE, Brier and reliability bins |
| Judge-score calibration | Shipped | LLMJudgeStrategy(calibrate=True, min_fit=20) keeps an IsotonicCalibrator per judge instance, pairs each raw judge score with the served target's observed quality/success and re-fits on every outcome; raw scores are kept in the rationale (raw=) once the mapping is active; calibrator state is persisted with state()/load() |
| Confidence exposed per strategy | Shipped | posterior variance surfaced for LinUCB and IRT |
Exit criterion: on examples/eval_dataset.jsonl and one public suite, learned signals beat the
static task table by a margin larger than the measured label-noise floor; osr eval reports ECE <= 0.10;
paraphrase robustness >= 0.9; core dependency count remains zero.
Measurement (published with 0.5.0, partially met). Command:
osr -t <targets> [--slm slm.json] eval <suite> --baselines --calibration --robustness --limit 1000.
The public suites are the human-preference battles relabelled onto three tiers (llm-small,
llm-mid, llm-frontier) by osr collect; the SLM bundle was trained on Arena-55k + RouteLLM
battles and the suites below are held out.
| Suite | Router (learned) | Declarative | Static task table | Random | Cheapest | Paraphrase | ECE raw | ECE held-out isotonic |
|---|---|---|---|---|---|---|---|---|
examples/eval_dataset.jsonl (n=30, 17 targets) | 0.833 | 0.833 | 0.233 | 0.133 | 0.067 | 0.867 | 0.245 | 0.261 (n=15) |
| MT-Bench human (n=1000) | 0.519 | 0.421 | 0.541 | 0.441 | 0.372 | 0.973 | 0.316 | 0.050 |
| PPE human (n=1000) | 0.561 | 0.418 | 0.651 | 0.445 | 0.174 | 0.898 | 0.279 | 0.036 |
| WebDev Arena (n=1000) | 0.533 | - | 0.670 | 0.391 | 0.000 | 0.872 | 0.253 | 0.060 |
- Met: learned signals beat the declarative router by +10 to +14 pp on every held-out suite and beat random / cheapest / best-prior everywhere; zero core dependencies; paraphrase robustness >= 0.9 on two of three public suites. With a calibrator fitted on half the suite, held-out ECE is 0.04-0.06 (isotonic) and 0.05-0.13 (temperature), i.e. the <= 0.10 target is met after fitting.
- Not met: on the public suites the static task table (which, with no task-type labels, degenerates to
"always the best single tier") still wins by 2-14 pp, because a pairwise preference between two
models is a weak label for a three-tier decision; the label-noise floor is not measurable on these
rows (one sample per prompt,
noise_floor()needssamples); raw, unfitted ECE is 0.25-0.32. - What this means: ship the router with a fitted
TemperatureScaler/IsotonicCalibrator(seecalibration_report()["held_out"]), and use multi-sample suites (RouterBench presets) to beat the task table by more than the noise floor - that measurement is tracked under v1.0 leaderboard runs.
Re-measurement (0.5.x SLM recipe, leave-one-source-out). The routing SLM now trains on ten public
sources, so each suite below was held out in turn and the SLM trained on the other nine (same command,
same targets, --limit 1000). The recipe change that matters for mixed catalogues - a corpus of model
battles no longer pushes down tools, skills and agents it never compared, and the strategy abstains on
them - takes examples/eval_dataset.jsonl with the SLM from 0.43 to 0.80 (0.83 without it; the old
recipe made any SLM harmful there). On the tier suites the numbers are unchanged within noise: PPE 0.561,
MT-Bench 0.508, WebDev Arena 0.506 (declarative 0.42 / 0.42 / -, static table 0.651 / 0.541 / 0.670);
held-out isotonic ECE 0.03 / 0.07 / 0.07. The conclusion stands: learned signals beat the declarative
router by +9 to +14 pp, and a pairwise battle remains too weak a label to beat "always the best tier".
v0.5: Target representations, effort and personalisation (Complete)#
Goal: learn what a target is good at from its own outcomes, so examples and quality_prior are
derived, not hand-written, and widen "target" to cover how hard the model should think and for whom.
0.5a Representations#
| Item | Status | Notes |
|---|---|---|
| Target embeddings from outcome logs | Shipped | learning.target_embedding(): centroid of description, capability lexicon and examples; nearest_targets() routes to unseen targets via neighbours (UniRoute / EmbedLLM) |
| Offline reward-matrix warm start | Shipped | learning.warm_start_from_matrix(rows, strategies, weight=): replays a full-information (prompt, {target: reward}) matrix into every learner before going online (LinUCB ridge per arm, IRT ability, Thompson posteriors, task table, Markov rewards); weight < 1 shrinks the offline evidence toward the prior (OrcaRouter 2605.30736) |
Automatic examples mining | Shipped | learning.ExampleMiner promotes high-quality outcomes into target.examples, de-duplicated by cosine |
| Cold-start via IRT warm-up | Shipped | learning.warm_start() copies shrunk IRT ability, Thompson posteriors and Bradley-Terry strengths from neighbours (IRT-Router) |
| Target similarity to fallback ordering | Shipped | learning.SimilarityFallback |
| Query-response mixed representation | Shipped | signals.DraftResponseSignal: hedges, self-corrections, query overlap and length ratio of a cheap draft (context["draft"] or a drafter callable) become extra["draft_uncertainty"] and lift complexity (JiSi 2601.01330) |
| Hidden-state Dirichlet router | Shipped | math.DirichletProbe (evidential Dirichlet head, digamma loss, KL regulariser, pure Python SGD) and strategies.HiddenStateStrategy(state_fn, targets, dim); epistemic uncertainty K/S becomes the strategy confidence so an unfamiliar state defers to the other strategies (ProbeDirichlet, RouterXBench 2602.11877) |
0.5b Effort as a routable dimension#
| Item | Status | Notes |
|---|---|---|
Target.effort variants | Shipped | RouteTarget.effort in {none, low, medium, high} with RouteTarget.family grouping siblings for breaker / budget purposes |
| Think-vs-non-think decision | Shipped | EffortStrategy / decide_effort() match ReasoningNeedSignal to effort level (Think When Needed 2601.18146; When to Think Deeply 2606.06745) |
| Token-budget-aware routing to elastic models | Shipped | strategies.expand_elastic(parent, [BudgetVariant(...)]) turns one elastic model into <id>@<budget> siblings sharing family and handler; TokenBudgetStrategy scores the fit between token_budget and the tokens the answer needs (truncation risk vs unused capacity) (Nemotron Elastic 2511.16664, Star 2605.07182) |
0.5c Personalisation#
| Item | Status | Notes |
|---|---|---|
RouteRequest.profile | Shipped | opaque user / tenant profile features; ProfileSignal hashes them into the LinUCB context |
| Few-shot user adaptation | Shipped | learning.UserAdaptiveStrategy: per-user x target Beta posteriors shrunk toward neighbour users (cosine over hashed profile vectors) and the global posterior, w = n / (n + kappa); observe_user() warm-starts from logged interactions; state()/load() (GMTRouter 2511.08590) |
| Counterfactual personalisation test | Shipped | eval.robustness.profile_swap_fairness() measures how often the decision changes with the profile (SkillFeed methodology 2608.28241) |
Exit criterion: adding a new target with only id, kind and cost reaches >= 90 % of its
hand-configured routing accuracy after 200 outcomes; effort routing reduces tokens >= 30 % at equal
quality on a reasoning benchmark. First half met: eval.criteria.cold_start_ratio() strips the
legal expert down to id/kind/cost next to seven configured experts and a generalist, streams
mixed traffic through the real Router with explore_rate=0.1, and measures 1.000 x the
hand-configured accuracy after 200 outcomes for the newcomer (0.000 before learning; 1 643 requests,
180 of them exploratory). Making this pass required cold-start exploration in the router
(Router(explore_rate=, explore_min_samples=)), because a target with no description or examples is
never ranked first and so never earns an outcome. Reproduce with python scripts/exit_criteria.py.
Effort-token half met in eval.criteria.effort_token_savings(): one reasoning model exposed through
expand_elastic() as reasoner@fast (effort="low") and reasoner@think (effort="high"), 2 000
rows carrying per-sibling scores and tokens (60 % plain lookups / rewrites where both siblings
score 0.90, 40 % proof / analysis prompts where the fast sibling drops to 0.50 - the pattern Think When
Needed and When to Think Deeply report for live thinking / non-thinking pairs). The real Router with
default_strategies() + EffortStrategy() + TokenBudgetStrategy() sends 59 % of the rows to the fast
sibling and spends 61.6 % of the always-think tokens (1 234 vs 2 003 per row; oracle 1 231) at
quality 0.900 vs 0.900. EvalRow.tokens is the new dataset field and osr eval DATASET --effort runs the
same comparison on rows collected from live thinking / non-thinking pairs; the synthetic figure is the
roadmap claim until a live pair is published from the hosted platform.
v0.6: Multi-step, multi-turn and agentic routing (Complete)#
Goal: treat routing as a sequence, not a single pick; align the unit of learning with the unit of supervision (the task).
0.6a Data model#
| Item | Status | Notes |
|---|---|---|
Outcome.task_id, Outcome.step, Outcome.role | Shipped | populated by Router.run(); learning.TaskCredit joins a terminal task-level reward to every call in the trajectory (TRACE-Router 2607.22465) |
| Delayed-feedback bandit updates | Shipped | TaskCredit buffers outcomes per task and credits on completion with uniform / discounted / last / blended schemes (Joulani et al. 2013) |
Session history in RouteRequest.history | Shipped | HistorySignal turn-level features; learning.HistoryTargetStrategy scores each target by a logistic model over the history-target joint embedding h * e_t (recency-weighted hashed history, catalogue target embedding, shared weights + per-target bias, online SGD), with an incumbent bonus that flips to a penalty when context["last_failed"] is set; state()/load()/merge() (MTRouter 2604.23530) |
0.6b Strategies#
| Item | Status | Notes |
|---|---|---|
| Admission-time task routing | Shipped | learning.TaskPins pins a task to its first target and releases on failure or after max_steps (TRACE-Router; MTRouter) |
| Cascade as MDP with a stop action | Shipped | CascadePlanner(mode="mdp"): finite-horizon expected value over candidate orderings (RLCascadeRouter 2608.15817; Dekoninck 2024). CascadeResult.outcomes() / Router.learn_cascade() turn every executed step into a per-step Outcome (rejected = failure with the gate score, accepted = success) that every learner consumes |
| POMDP self-verification cascade (AutoMix) | Shipped | CascadePlanner(mode="pomdp") with BeliefTracker |
| Bayesian self-escalation | Shipped | strategies.SelfEscalation keeps a Beta competence posterior updated per streamed chunk (hedges, refusals, self-corrections, repetition, overrun) and stops when loss(stop) < loss(continue); wrap_stream(chunks, monitor) cuts any text iterator at the escalation point; SelfEscalation.for_target() seeds the prior from quality_prior (2608.24087) |
| Permanent-handoff policy from censored teacher signals | Shipped | learning.MixtureCureModel (Weibull mixture-cure on cumulative risk, grid MLE with censoring) and learning.HandoffPolicy(threshold, hazard_threshold, horizon_steps, fallback_target, pins): hands a task over once and releases its TaskPins pin (TACIT-Switch 2608.27911) |
| Progress-guided step routing | Shipped | ProgressRouter: step index, failures, budget spent and last target condition each step's route (ProgRouter 2608.25992; 2511.06190) |
| Routing / aggregation switch | Shipped | strategies.MixtureOfAgents(k, confidence_below, budget_usd, aggregator): calls the top-k alternatives when the decision is unsure and the budget allows, aggregates by semantic majority vote (majority_vote) or a supplied synthesiser, and emits one Outcome per participant (JiSi 2601.01330; Mixture-of-Agents) |
| Multi-round executor (Router-R1) | Shipped | MultiRoundExecutor: route, execute, judge, refine loop with bounded rounds |
| Plan execution | Shipped | Router.run() / execution.py: skill pre-processes, persona and skill instructions set the system prompt, primary answers; per-participant outcomes; slot_quality_floor |
| Conversation-state MDP from real sessions | Shipped | MarkovStrategy(state_from="auto") keys transitions and rewards on the learned task type (ontology leaf from the task-type classifier) joined with the domain, "task_type" / "domain" force one axis; per-request state memory so the outcome credits the state that was scored; decay= forgets stale sessions |
| Collaboration-protocol selection | Shipped | strategies.ProtocolPolicy: failure_risk() from decision confidence, signals and response uncertainty; an ordered ProtocolRule table (risk, budget, task type) picks single / cascade / aggregate / debate / handoff; per-protocol ledger reports which protocol actually paid (2608.14927) |
Exit criterion: on tau-bench style agentic tasks, task-level routing beats the best single target on
the accuracy-latency frontier; on multi-hop QA the executor beats the best single target at <= 60 % of
its cost, reproducibly via osr eval --multi-round; cascade outcomes are consumed by every learner.
Multi-round half met on the synthetic scored dataset in eval.criteria.multi_round_vs_best_single()
(1 000 rows, small / medium / large targets with per-row scores): MultiRoundExecutor with
failures_before_switch=1 reaches quality 0.911 (best single target large: 0.900, oracle 0.912) at
49.8 % of the best single target's cost, 1.8 rounds on average. osr eval DATASET --multi-round
runs the same comparison on any dataset whose rows carry scores. Cascade outcomes are consumed by
every learner (Router.learn_cascade(), tested).
tau-bench half met in eval.agentic.task_routing_frontier(): 1 000 synthetic tasks of 3-5 steps
(customer support, coding, travel, finance; 30 % reasoning-heavy steps) with a per-step success /
latency table for three agents - fast-agent (400 ms, 0.97 on plain steps, 0.45 on hard ones),
balanced-agent (900 ms, 0.96 / 0.72) and strong-agent (2 000 ms, 0.94 / 0.90) - so no single agent
dominates per step, the shape tau-bench reports across model tiers. Each agent is a CallableHarness
behind harness_handler(); the real ProgressRouter (escalate on failure, switch after two) routes every
step with EffortStrategy and the auto-learning quartet, TaskCredit credits the task result back, and
both policies get one retry per step on common random numbers. Routed accuracy 0.987 vs 0.977 for the
best single agent (strong-agent) at 45.4 % of its latency (3.9 s vs 8.5 s per task; 67 % of steps
go to fast-agent, 33 % to strong-agent). AgentTask / load_agentic_tasks() define the JSONL
format and osr eval TASKS.jsonl --agentic runs the same comparison on a table collected from live
harnesses; the synthetic frontier is the roadmap claim until a live one is published.
v0.7: Catalogue interop and discovery at scale (Complete)#
Goal: stop hand-writing targets.yaml for things that already describe themselves, and keep routing
accurate when the catalogue has thousands of entries.
0.7a Import#
| Item | Status | Notes |
|---|---|---|
MCP tool catalogue import to TargetKind.TOOL | Shipped | adapters.mcp.tools_from_mcp() (offline tools/list), connect_mcp() (stdio JSON-RPC), enrich_description() at import time (Enrich-Retrieve-Rank 2608.22695) |
| MCP server recommendation | Shipped | adapters.recommend_servers(task, servers, k, constraints, allowed_auth): BM25 + hashing-cosine fusion over ServerCard text (name, description, tool names, tags) with hard filters on latency, region, data boundary and auth kind; rationale names the matched tools (Task2MCP / T2MRec 2604.17234; ComplexMCP 2605.10787) |
A2A agent cards to TargetKind.AGENT | Shipped | adapters.a2a.fetch_agent_card(), agent_from_card(), skills_from_card(), a2a_handler() |
Agent-harness adapters to TargetKind.AGENT | Shipped | CallableHarness, HTTPHarness, SubprocessHarness with a common HarnessResult and task_id-carrying outcomes (2607.11399; SWE-Router 2607.00053) |
Skill packages to TargetKind.SKILL | Shipped | load_skills() reads agentskills.io SKILL.md folders |
| Persona catalogue import | Shipped | adapters.personas.load_personas() from Markdown, JSON, CSV and Copilot agent / chatmode files; persona choice learned from plan-level outcomes (MasRouter) |
| OpenAI-compatible proxy mode | Shipped | POST /v1/chat/completions with model: "auto" routes, executes and learns; GET /v1/models |
| LangGraph / Microsoft Agent Framework node | Shipped | adapters.frameworks.langgraph_node(), langgraph_condition(), maf_router_executor() |
| vLLM Semantic Router / gateway parity | Shipped | adapters.load_semantic_router_config(path_or_dict) imports a semantic-router config.yaml (categories, model_scores, use_reasoning, model_config.pricing, system_prompt) as RouteTargets plus category Rules, so an existing gateway config routes through OpenSmartRoute unchanged (2603.04444) |
0.7b Two-stage selection#
| Item | Status | Notes |
|---|---|---|
| Retrieve-then-rank narrowing | Shipped | retrieval.Retriever: BM25 + hashing-dense retrieval with reciprocal-rank fusion, on above narrow_above (default 500) targets (Enrich-Retrieve-Rank; SCOUT 2608.23992) |
tool_search / execute_tool meta-tools | Shipped | retrieval.tool_search_target() and execute_tool_target() expose the router as two MCP tools (SCOUT) |
| Field-aware / schema-aware tool matching | Shipped | discovery.schema_match(request, target) extracts typed entities (emails, URLs, dates, times, paths, ids, numbers, currency, quoted strings) and scores coverage of the tool's required input_schema properties; discovery.SchemaAwareStrategy fuses it with the other strategies (SchemaRouter 2608.21375) |
| Cache-preserving tool routing | Shipped | discovery.CachePreservingSelector(prefix_size, evict_after) keeps a per-session ordered tool prefix byte-stable across turns, appends new tools and evicts only after inactivity; prefix_hit_ratio() reports cache stability (CacheRouter 2608.22708) |
| Submodular skill-set selection | Shipped | retrieval.select_skill_set(): greedy benefit minus redundancy under a token budget (Best Prefix Selection 2608.19993) |
| Profile-conditioned skill routing | Shipped | learning.SkillAffinity: Beta posterior per (profile bucket, skill); relevance(profile, skills) plugs into select_skill_set(relevance=) (SkillFeed 2608.28241) |
| Skill graphs / composition | Shipped | discovery.SkillGraph: requires / conflicts / composes edges from SKILL.md frontmatter (osr-requires, osr-conflicts, osr-composes) plus learned co-usage; compose(seeds) closes under dependencies, resolves conflicts by priority and returns a dependency order (2606.18051; CaSKG 2608.25500) |
0.7c Shared state#
| Item | Status | Notes |
|---|---|---|
Redis / Postgres StateStore | Shipped | enterprise.stores.RedisStateStore, SQLStateStore (DB-API 2.0), BatchedStateStore for write coalescing, NamespacedStateStore |
| State schema version and migration | Shipped | VersionedStateStore stamps {"_schema": n} and runs forward migrations; refuses to load newer state |
Exit criterion: a fresh install routes to an MCP server's tools with no YAML written; with 5 000
synthetic tools Match@1 stays within 5 points of the 50-tool figure. Met.
eval.criteria.match_at_1_at_scale() builds 50 and 5 000 verb-object-qualifier tools
(synthetic_tool_catalogue()), routes 200 paraphrased prompts through RouterBuilder.with_retrieval()
and measures Match@1 1.000 -> 1.000 (drop 0.000).
v0.8: Operations, risk control and economics (Complete)#
Goal: make the learners safe to leave running: bounded risk, bounded budgets, honest under drift, measurable without live traffic.
0.8a Risk control and abstention#
| Item | Status | Notes |
|---|---|---|
| Conformal escalation thresholds | Shipped | math.calibration.ConformalCalibrator (split conformal) behind RouterBuilder.with_calibration() (RouteNLP 2604.23577) |
| Set-valued routing with abstention | Shipped | RouteDecision.candidate_set = smallest nested set with mis-routing risk <= alpha (RACER 2603.06616; CR2 2605.12001) |
Learning-to-defer for TargetKind.HUMAN | Shipped | DeferStrategy (Madras 2018; Mozannar and Sontag 2020; Verma 2023) |
| Uncertainty features from the response | Shipped | signals.semantic_entropy() / response_uncertainty() (meaning clusters over sampled answers, hedging, optional P(True) judge); signals.UncertaintyGate is a drop-in Cascade quality gate and MixtureOfAgents trigger; signals.EventTrigger fires named actions from threshold rules (Kuhn 2023, Kadavath 2022; 2607.13048) |
0.8b Non-stationarity and budgets#
| Item | Status | Notes |
|---|---|---|
| Forgetting in every learner | Shipped | decay < 1 for ThompsonBeta, IRTModel, BradleyTerry, MarkovChain (row counts) and RoutingMDP (exponentially weighted rewards); discount < 1 for LinUCB with an undiscounted ridge floor (D-UCB, Garivier and Moulines 2011); MarkovStrategy(decay=) wires both Markov knobs |
Multi-knapsack CostAwareBandit | Shipped | math.bandits.MultiKnapsackBandit: several resources (USD, tokens, GPU-seconds, energy), one shadow price per resource by dual ascent, hard sliding-window meters checked before commitment, learned optimistic costs, pessimistic rewards, bounded audit stream of every pruned arm (Drift-Aware Sparse Routing 2609.00662; Badanidiyuru 2013) |
| Tenant fairness under shared budgets | Shipped | enterprise.ops.FairShareMiddleware: dominant-resource fairness over a sliding window |
| Numerically stable LinUCB | Shipped | Sherman-Morrison rank-1 updates; surfaced posterior variance |
0.8c Latency, hardware and energy#
| Item | Status | Notes |
|---|---|---|
| Queue-aware latency estimate | Shipped | enterprise.ops.InflightTracker + QueueAwareStrategy (Erlang-C / Kingman) (Latency-aware routing 2607.18253) |
| Energy / carbon as cost dimensions | Shipped | RouteTarget.cost["wh_per_1k_tokens"] / ["gco2_per_1k_tokens"] (or gco2_per_wh grid intensity) read by unit_energy / unit_carbon; Objective(energy=, carbon=) weights them in the utility, log-normalised like money; MultiKnapsackBandit meters them as resources |
| Hardware-aware cost | Shipped | math.EnergyModel fits Wh = e0 + e_in * prompt + e_out * output per target by ridge regression from measured samples and converts to gCO2 with a grid factor; math.HardwareProfile / hardware_profile() give priors for A100, H100, L4, RTX 4090, CPU and NPU when no meter exists; feeds wh_per_1k_tokens / gco2_per_1k_tokens (2608.28044; HW-Router 2608.14575) |
| Edge-cloud token-aware routing | Shipped | strategies.EdgeCloudStrategy: targets tagged `metadata["tier"] = edge |
0.8d Offline and online evaluation#
| Item | Status | Notes |
|---|---|---|
| Off-policy evaluation | Shipped | eval.ope: IPS / SNIPS / doubly-robust with weight clipping and effective sample size; osr ope (Dudik et al. 2011); RouteDecision.propensities logs the softmax routing policy so LoggedDecision.from_decision needs no extra plumbing |
| Shadow mode and A/B harness | Shipped | enterprise.ops.ShadowMiddleware and ABTest with Wald's SPRT for promotion |
| Federated learner merge | Shipped | merge() on ThompsonBeta, LinUCB, IRTModel, BradleyTerry, MarkovChain / RoutingMDP, BanditStrategy and TaskTableStrategy combine sufficient statistics without sharing prompts; learning.merge_learners(local, remote) pairs same-named strategies across replicas ("Federate the Router" 2601.22318). AutoLearner.refresh() is the writer / reader pattern for replicas that share one StateStore |
| Market / auction strategy | Shipped | strategies.AuctionStrategy: targets bid (claimed success, price); claims are corrected by each bidder's observed bias and realised rate, the highest corrected surplus wins and pays the second price (EA-RAM 2608.12719) |
Exit criterion: conformal thresholds hit their nominal mis-routing rate +/- 2 pp on held-out data;
the knapsack bandit never exceeds a hard budget in a 10^6-step simulation with drift; IPS estimates of a
held-out strategy are within the bootstrap CI of its live result. Met (python scripts/exit_criteria.py --full):
conformal_coverage():ConformalCalibrator(alpha=0.1)fitted on 1 000 router propensities over graded mixed-domain prompts covers 0.909 of 1 000 held-out labels (gap 0.9 pp, mean set size 1.2).knapsack_never_exceeds_cap():MultiKnapsackBandit(on_capped="abstain")over 10^6 steps with reward means reshuffled every 10^5 steps; peak sliding-window spend / cap = 0.9998 (USD) and 0.9996 (tokens), zero abstentions, mean reward 0.75. Getting there fixed two real defects: the "release the cheapest arm" fallback could breach the cap (peak 1.0165), and shadow prices exploded for token-scale budgets.ope_within_live_ci(): 2 000 decisions logged under a cost-heavy softmax policy; IPS 0.812, SNIPS 0.816 and DR 0.816 all fall inside the live value's 95 % bootstrap CI [0.798, 0.832] (live 0.815, n_eff 992).
v0.9: Security hardening of the control plane (Complete)#
Goal: close the gaps listed in SECURITY.md that heuristics cannot.
| Item | Status | Notes |
|---|---|---|
| Learned gadget / confounder detector | Shipped | security.gadget.GadgetDetector; InputGuard(learned=True); trained via osr train --gadget (Rerouting LLM Routers) |
| Origin policy for tool parameters | Shipped | security.provenance.OriginPolicy: sensitive parameters of state-changing tools must originate from the user turn (ROPE 2608.27496) |
| Signed MCP manifests | Shipped | adapters.mcp.sign_manifest() / verify_manifest() (HMAC-SHA256 or Ed25519); osr mcp-manifest (2601.23132) |
| Resource-amplification limits | Shipped | security.limits.ResourceLimiter: per-task caps on steps, tool calls, depth, tokens, cost and wall-clock (Beyond Max Tokens 2601.10955) |
| Safety-routing regression suite | Shipped | security.safety.run_safety_suite(); osr safety --learned-guard gates CI (When Safety Routing Breaks 2609.01455) |
| Encrypted state at rest | Shipped | enterprise.stores.EncryptedStateStore (AES-256-GCM, key rotation) |
| External review pack | Complete | SECURITY_REVIEW.md: scope, trust boundaries, evidence table with reproduction commands, reviewer questions, known gaps, review log |
Exit criterion: red-team suite passes in CI (met); published threat-model delta (met: SECURITY.md, "Threat-model delta: 0.3 -> 0.4"); a published review pack that lets an independent party run the review without a briefing (met: SECURITY_REVIEW.md).
The original wording also required the independent report itself. That cannot be produced from inside
the repository, so it is now the adoption item it really is: tracked in the v1.0 readiness table below
and closed when a third-party report is added to the review log in SECURITY_REVIEW.md.
v1.0: Stable API (Complete)#
Goal: freeze the public surface.
- All
opensmartroute.*public symbols frozen; the snapshot test (tests/public_api.json) already fails on any added or removed name. Deprecation policy: shipped in CONTRIBUTING.md, "Public API and deprecation policy" witherrors.deprecated()/OpenSmartRouteDeprecationWarningas the mechanism. - Documented performance envelope: shipped in ARCHITECTURE.md
(p50 / p95 / p99 for 16 / 64 / 256 / 1024 targets, with and without two-stage narrowing;
python scripts/bench.py --scalereproduces it). - Security review of the control plane against
docs/SECURITY.mdthreat model: the review pack is shipped (SECURITY_REVIEW.md); the independent report is an adoption item (table below). - Reference deployments: shipped as library (
pip install), sidecar (deploy/Dockerfile,osr serve) and central control plane (deploy/helm/opensmartroute; hardened pod security, HPA, PDB, NetworkPolicy, ConfigMap-driven catalogue) - see deploy/README.md; hosted platform on Azure Container Apps with Azure OpenAI (platform/,infra/,azd up; community and enterprise editions, API keys, plans, metering, tenants, audit) - see platform/README.md. - Public leaderboard runs with reproducible configs: shipped - examples/leaderboard/results holds the 0.5.0 run (learned router vs declarative, static task table and baselines on MT-Bench human, PPE human and WebDev Arena, with robustness and held-out calibration; suite-level hold-out, seed 0, SLM SHA-256 recorded) produced by the recipe in examples/leaderboard. Listing on RouterArena, LLMRouterBench and xRouteBench is a submission to those projects and is tracked below.
Exit criterion (settled with 1.0.0). 1.0 means what the repository can prove: a frozen public API with a
deprecation path, a published performance envelope, reference deployments, the security controls with a
red-team suite in CI and a review pack, a published leaderboard run, and two consecutive releases without a
breaking change. Adoption evidence - an independent security report, accepted leaderboard listings, two
production users - needs a third party, so it is tracked and reported but does not gate the version.
Nothing is typed by hand: python scripts/release.py readiness computes the table below from the
repository (snapshot and tests present, envelope section, deployment files, review pack and CI safety
suite, published run, non-breaking release steps in CHANGELOG.md verified against the tagged
tests/public_api.json, review-log rows in SECURITY_REVIEW.md, accepted rows under Listings in
the results page, rows in ADOPTERS.md); release.py check refuses any 1.x version
while a release row is open, and the Release workflow prints the whole table in its summary.
| Readiness item | Kind | How it is verified | Status |
|---|---|---|---|
| Public API frozen with a deprecation path | release | tests/test_public_api.py fails on any removed or renamed name; errors.deprecated() | Met |
| Performance envelope published | release | ARCHITECTURE.md, scripts/bench.py --scale | Met |
| Reference deployments (library, sidecar, control plane, hosted) | release | deploy/, platform/, live at opensmartroute.ai | Met |
| Security controls, red-team suite in CI, threat-model delta, review pack | release | v0.9 above | Met |
| Leaderboard run published with reproducible config | release | examples/leaderboard/results | Met |
| Two consecutive releases with no breaking change | release | 0.4.0 -> 0.5.0 and 0.5.0 -> 1.0.0, additions only against the tagged tests/public_api.json | Met |
| Independent security report | adoption | A row in SECURITY_REVIEW.md, "Review log" | Open |
| Leaderboard listings accepted | adoption | An accepted row under results, "Listings" linking the published run | Open |
| Two independent production users | adoption | Two rows in ADOPTERS.md (named, or anonymised with consent) | 0 of 2 |
What closes each adoption row:
- Security report - a third party runs SECURITY_REVIEW.md and adds a row to its review log with the report link; findings are fixed or accepted first.
- Listings - submit the 0.5.0 run (commit, SLM SHA-256, results page) to RouterArena,
LLMRouterBench and xRouteBench; record the entry URL and
acceptedin the results page. - Production users - two organisations agree to be listed in
ADOPTERS.md.
After 1.0 the public API changes only through the deprecation policy: a removal needs a
deprecated() warning for at least one minor and a ### Removed entry, which the readiness check
reads as a breaking step and the snapshot test refuses without an explicit snapshot update.
Research track#
Every question below now has a runnable answer in the tree; the table records where, so the question can be re-asked against a real catalogue and dataset rather than argued about.
| Question | Answer in the code | Related work |
|---|---|---|
| Is the remaining router-to-oracle gap real or label noise on our pools? | eval.headroom.routing_headroom(rows, targets): oracle vs best single target, split into the label-noise floor (noise_floor) and the measurable, recoverable headroom | 2607.03436; LLMRouterBench 2601.07206 |
| Does a static task table already capture most of the value for typical enterprise catalogues? | eval.headroom.learnability_by_difficulty buckets rows by oracle-minus-best-single gap; TaskTableStrategy is a mandatory osr eval --baselines row, so the table's share of the headroom is reported per run | 2608.23023 |
| How few, and how different, must targets be before routing beats the best single model? | eval.headroom.target_diversity (mean pairwise disagreement, winner entropy / share) and min_catalogue(rows, targets, fraction) (smallest subset keeping 95 % of the oracle) | 2607.09197 |
| Contrastive router training on (query, target) pairs | learning.ContrastiveRouter (softmax cross-entropy over q·e_t / τ against the acceptable set of a row, acceptable_set(row, slack)) and ContrastiveStrategy with online updates from outcomes | RouterDC (NeurIPS 2024); RouterDC-style objectives in LLMRouter |
| Reward-model-distilled routing labels | learning.soft_labels(scores, temperature) and ContrastiveRouter.fit(objective="distilled"): train against the softened reward distribution instead of one-hot labels | Zooter |
| Can the whole ensemble be compressed into one small routing model? | learning.RouterSLM (hashed dual encoder + temperature calibration + target snapshots in one JSON file) and distill_router(router, texts), which trains it on the router's own propensities; SLMStrategy puts the student back in the ensemble and keeps learning online; osr slm train/eval/predict | RouteLLM matrix-factorisation router; Zooter; RouterDC |
| Does a router keep improving on its own from live data, without regressing? | learning.SelfImprover: champion/challenger cycle over feedback rows (eval.rows_from_feedback), Hub datasets (eval.DatasetCollector, RouterBench / RouterEval) and synthetic rows; challenger is calibrated on the holdout and promoted only above OSR_SLM_MIN_GAIN; JSONL history; osr improve | RouterBench 2403.12031; RouterEval 2503.10657 |
| Where do current prices and quality priors come from at scale? | adapters.ModelCatalogue (OpenRouter prices, Hugging Face model-index benchmarks and popularity → quality_prior, Pareto frontier(), cost_benchmark()) and adapters.WebKnowledge (Hugging Face / DuckDuckGo / Brave discovery, risk-scored page text); osr catalogue | LLMRouterBench 2601.07206 cost-quality frontier |
| Model-level scaling: does accuracy keep rising with more targets? | eval.headroom.scaling_curve(rows, targets, sizes): oracle / best-single / headroom for random sub-catalogues of each size, exposing diminishing returns | RouterEval; LLMRouterBench diminishing returns |
| Intra-generation escalation vs pre-generation routing | strategies.SelfEscalation + wrap_stream (stop mid-stream when loss(stop) < loss(continue)); compare against Cascade on the same dataset | 2608.24087 |
| Routing as an RL problem end-to-end (Router-R1, RLCascadeRouter) vs predict-then-optimise | learning.PolicyGradientStrategy (REINFORCE on hashed request features, EWMA baseline, entropy bonus; reward is the decision utility decision_reward) and learning.decision_regret(rows, targets, choose, predict) which reports decision regret next to prediction MAE so the two losses can be compared directly | 2608.15817 |
| Is routing learnable where it is most valuable (web / SWE agents)? | eval.headroom.learnability_by_difficulty(rows, targets, buckets): per-difficulty-bucket oracle, best single, headroom and noise floor; the hardest bucket answers the question for a given pool | 2608.06171; SWE-Router 2607.00053; Agent-as-a-Router 2606.22902 |
| Memory-tier routing for agents | strategies.MemoryRouter: tiers with capacity, write/read cost and latency; learned ImportanceGate decides what to write, value decays per tick, lowest-value residents are demoted, recall(query, budget_tokens) fills a token budget by relevance | BudgetMem 2602.06025; Gated-Memory Routing 2609.00237 |
| Multimodal / modality escalation | strategies.ModalityStrategy (penalises text-only targets when the request carries images / audio / video / files that only a text surrogate describes) and ModalityEscalation (text-first with a Beta posterior on "surrogate sufficed", escalates to the multimodal target when confidence is low) | LatentRouter 2605.11301; modality escalation 2607.05438; CUA routing 2603.12823 |
| Mechanism design when providers are strategic | strategies.AuctionStrategy: bias-corrected claims, second-price payment | EA-RAM 2608.12719; strategy auctions 2602.02751; coalition pricing 2608.07532 |
| Semantic caching as a target | strategies.SemanticCache.target() is a DESTINATION whose handler serves near-duplicate answers; SemanticCacheStrategy scores it from the hit similarity and learns per-band hit quality from outcomes | GPTCache; vLLM semantic router 2603.04444 |
| Routing among human annotators / experts | strategies.AnnotatorPool (Dawid-Skene EM over per-domain annotator accuracy), select_quorum(domain, target_accuracy, budget) (greedy majority-vote quorum via quorum_accuracy) and HumanRoutingStrategy for HUMAN targets | QUORUM 2608.27974 |
| Speculative / draft-based cascades | strategies.SpeculativeCascade: runs draft and strong in parallel when the learned acceptance posterior says speculation is cheaper than either alone (expected_mode_costs), cancels the strong call when the draft is accepted | speculative cascades; Differential Reasoning Router 2608.30224 |
Open follow-ups (not scheduled): learned energy/carbon priors per provider region, and replacing the
hashed encoders in ContrastiveRouter / PolicyGradientStrategy with a real embedding model behind
Embedder when a dependency is acceptable.
Out of scope#
- Model serving, quantisation, or inference optimisation (use vLLM / TGI).
- Protocol translation between provider APIs (use LiteLLM / aisix in front).
- Token-level / layer-level routing inside a model (MoE, early exit, hybrid attention).
- Prompt management / versioning.
- Being a chat UI.