<!-- OpenSmartRoute: osr-enterprise-builder. Source https://github.com/isathish/OpenSmartRoute/blob/main/.claude/skills/osr-enterprise-builder/SKILL.md; HTML https://opensmartroute.ai/docs/skills/osr-enterprise-builder -->
---
name: osr-enterprise-builder
description: Assemble a production OpenSmartRoute router with RouterBuilder - auto-learning (IRT, Bradley-Terry, LinUCB, Markov), health circuit breakers and budgets, tenant/cache/timeout/guard middleware, logging and metrics telemetry, hash-chained audit, Redis/SQL/encrypted state stores, calibration, retrieval narrowing, shadow/A-B rollout and fair share. Use when hardening a router for multi-tenant or high-volume use, persisting learner state, or adding observability.
license: Apache-2.0
compatibility: OpenSmartRoute >= 0.4, Python >= 3.10
metadata:
  author: opensmartroute
  osr-domains: "coding general"
  osr-tags: "opensmartroute enterprise builder middleware telemetry"
  osr-quality-prior: "0.85"
  osr-primary: "false"
---

# OpenSmartRoute enterprise builder

`RouterBuilder(registry, settings=None)` is a fluent builder; every `with_*` returns the builder and
`build()` returns an `EnterpriseRouter` (thread-safe facade: `route`, `learn`, `execute`, `run`,
`health_snapshot`). It raises `ConfigurationError` for an empty registry, no strategies, or duplicate
strategy names.

## Reference chain

```python
from opensmartroute.enterprise import (RouterBuilder, CacheMiddleware, TenantMiddleware, TimeoutMiddleware,
                                       LoggingTelemetry, MetricsTelemetry, FileAuditSink)
from opensmartroute.observability import MemorySink, MetricsSink
from opensmartroute.security import GuardMiddleware

metrics = MetricsTelemetry()
app = (
    RouterBuilder(registry)
    .with_rules(load_rules("rules.yaml"))            # declarative preferences (RulesStrategy)
    .with_defaults()                                 # capability + similarity + Thompson bandit
    .with_auto_learning(state_dir=".osr-state")      # IRT + Bradley-Terry + LinUCB + Markov, persisted
    .with_health(latency_slo_ms=3000)                # breakers, rate limits, budgets; HealthPolicy + HealthStrategy
    .with_middleware(GuardMiddleware(redact=True),   # order = execution order (outermost first)
                     TenantMiddleware({"acme": {"max_cost_per_1k": 0.005, "data_boundary": "private",
                                                "deny_targets": ["llm-frontier"]}}),
                     CacheMiddleware(max_size=4096, ttl_s=30),
                     TimeoutMiddleware(budget_ms=250))
    .with_telemetry(LoggingTelemetry(), metrics)     # on_decision / on_outcome / on_error
    .with_tracing(MemorySink(), MetricsSink())       # spans + events of every stage (docs/OBSERVABILITY.md)
    .with_audit(FileAuditSink("audit.jsonl"))        # hash-chained, tamper-evident
    .with_calibration(conformal_alpha=0.1)           # temperature scaling + conformal candidate sets
    .with_retrieval(narrow_above=500, narrow_to=50)  # BM25 + dense RRF narrowing for big catalogues
    .with_llm_judge(LLMJudgeStrategy(judge), escalate_below=0.6)
    .with_objective(Objective(cost=0.3))
    .build()
)
```

Other options: `with_strategy(s, weight)`, `with_policy(policy)` (wrapped by `HealthPolicy` when
health is on), `with_feedback(FeedbackStore)`, `with_state_store(store, save_every)`,
`with_queue_awareness(slo_ms)` (in-flight tracking + Erlang-C latency), `with_fair_share(weights)`,
`with_shadow(candidate_router, mode="shadow"|"ab")` (SPRT-judged), `with_router_options(**kw)`,
`with_components(registry)` (decorator SDK).

Optional learners and controls added with `with_strategy(...)` (each is a `Strategy`; those with a
`model` are persisted by `with_auto_learning` / `with_state_store` like the built-in ones):
`learning.HistoryTargetStrategy()` (multi-turn: history-target joint embeddings),
`learning.UserAdaptiveStrategy()` (per-user shrinkage on `context["user_id"]`),
`strategies.EdgeCloudStrategy()` (`metadata.tier` edge/cloud vs latency SLO and data boundary),
`strategies.TokenBudgetStrategy()` (elastic `<id>@<budget>` siblings), `strategies.AuctionStrategy()`
(bias-corrected second-price bids), `strategies.HiddenStateStrategy(state_fn, targets, dim)`,
`discovery.SchemaAwareStrategy()` (tool `input_schema` coverage). After the answer:
`signals.UncertaintyGate` as the `Cascade` quality gate, `strategies.ProtocolPolicy` to pick
single / cascade / aggregate / debate / handoff per request, `strategies.MixtureOfAgents` for top-k
aggregation, `strategies.SelfEscalation` + `wrap_stream` for mid-stream escalation and
`learning.HandoffPolicy` for permanent hand-off of failing agent tasks. Energy and carbon:
`math.EnergyModel` / `hardware_profile()` fill `cost.wh_per_1k_tokens` / `gco2_per_1k_tokens` that
`Objective(energy=, carbon=)` weighs.

## Middleware contract

```python
class Middleware(ABC):
    def __call__(self, request: RouteRequest, next_: RouteFn) -> RouteDecision: ...
```
Raise `SecurityError`/`ValidationError` to reject; mutate `request.constraints` to enforce tenant
defaults (that is what `TenantMiddleware` does); wrap `next_(request)` for caching/timing.

## Telemetry contract

`Telemetry.on_decision(request, decision)`, `on_outcome(outcome)`, `on_error(request, error)`. Never
log raw text - `LoggingTelemetry` logs a SHA-256 prefix + length. `MetricsTelemetry.snapshot()` gives
`decisions_by_target`, `route_latency_ms{p50,p95,p99}`, `policy_rejections_total`, `mean_confidence`.
`opensmartroute.adapters.OpenTelemetryTelemetry` needs the `otel` extra.

## Tracing contract

Telemetry sees decisions; the tracer sees *how* they were made. `with_tracing(*sinks_or_tracer,
sample_rate=)` (no arguments: `Tracer.from_settings()` / `OSR_OBSERVABILITY_*`) records a `request` span per
call with nested `route` / `plan` / `execute` spans and events `route.signals`, `route.policy`, `route.rank`
(per-strategy ms), `cache.hit|miss`, `tenant.rejected`, `route.slow`, `guard.*`, `health.breaker`,
`shadow.*`, `fairshare.*`, `learn.outcome`. Sinks: `MemorySink` (`trace(request_id)`), `MetricsSink`
(`prometheus()`), `LoggingSink`, `FileSink`, `adapters.optional.OpenTelemetrySink`. Inside your own
middleware emit with `current_tracer().event("name", **attrs)` - never pass text, only ids and numbers.
See `docs/OBSERVABILITY.md`.

## Persisting learner state

`with_auto_learning(state_dir)` writes JSON files; for shared infrastructure use
`enterprise.stores`: `RedisStateStore`, `SQLStateStore`, plus wrappers `EncryptedStateStore`
(needs `crypto` extra), `VersionedStateStore(Migration...)`, `BatchedStateStore`,
`NamespacedStateStore`. Pass with `with_state_store(store, save_every=10)`. Multi-replica: one
writer, readers reload; state is quarantined on load failure rather than trusted blindly.

## Closing the loop

```python
d = app.route(RouteRequest("...", constraints=RequestConstraints(tenant="acme")))
app.learn(Outcome(d.request_id, d.target.id, success=True, quality=0.9, cost_usd=0.002, latency_ms=1200,
                  domains=d.trace.signals.domains, complexity=d.trace.signals.complexity))
res = app.run("...")          # route -> execute plan -> learn, with the full middleware chain
```

`app.health_snapshot()` exposes breaker state per `target.health_key` (family-aware);
`app.shadow` returns the `ShadowMiddleware` for A/B statistics.

## Checklist before shipping

- Middleware order: guard -> tenant -> cache -> timeout.
- `TenantMiddleware(require=True)` unless anonymous traffic is expected.
- Set `latency_slo_ms`, `Objective.quality_floor`, and per-tenant `max_cost_per_1k`.
- Persist state (`state_dir` or a `StateStore`) and mount it in the container (`OSR_STATE`).
- Export `metrics.snapshot()` and keep the audit sink on durable storage.
