PROJECT 04 / 18AI INFRASTRUCTUREGO

Routing gateway

ModelRouter.

Spend reasoning where it changes the answer.

Model × effortrouting decision unit
3 providersnative adapter families
Pure Golocal prediction path
01 / IDEA02 / SYSTEM03 / PLAYGROUND04 / DECISIONS05 / SOURCE
01 / THE IDEA

A closer look.

An OpenAI-compatible gateway that chooses a model and reasoning effort for each request. A local predictor estimates task difficulty and skill requirements; a constrained optimizer balances expected success, cost, and latency before a streaming dispatcher handles provider execution and failover.

A model name is a coarse allocation decision. ModelRouter treats model and effort together as an arm, estimates how well that arm fits the request, and makes the cost of a stronger answer explicit while preserving protocol compatibility for clients.

01

Predict locally

Structural request features combine with a pure-Go Model2Vec embedder. A nearest-neighbor exemplar bank and ridge head estimate difficulty, skill mix, output length, and uncertainty without calling a routing LLM.

02

Choose model and effort together

The optimizer filters model candidates by capabilities, credentials, and health, then evaluates each supported effort. Utility combines predicted success with cost and latency penalties; session state influences switching and cache-aware cost.

03

Keep failover before commitment

Upstream responses are streamed into a peek buffer. The dispatcher can switch attempts before the first visible token commits a stream, with effort-scaled timeouts, retry budgets, and optional hedging.

04

Close the feedback loop

Feedback and observed response signals update per-skill model offsets and exemplars. Spend control adjusts the cost penalty, while health tracking informs later routing decisions.

02 / UNDER THE SURFACE

Inside one routing decision

Local selection is separate from upstream execution; feedback reconnects the two.

DRAG TO PAN · SELECT A NODE · + / − TO ZOOM

Read the architecture as text
  1. Compatible request — The gateway accepts chat requests and a route-preview endpoint. Canonical request types preserve messages, tools, attachments, and per-request router preferences across provider-specific adapters.
  2. Structural features — Extract walks the canonical request to estimate input size and structural demand. These features complement semantic embeddings: tools, schema requirements, attachments, and code all affect routing suitability.
  3. Static text embedding — The embedding implementation tokenizes text and pools static token vectors locally. It returns the semantic representation used by the predictor without an extra network inference call.
  4. Difficulty + skill mix — predictVec weights nearby labeled exemplars, blends them with ridge predictions, and estimates uncertainty from disagreement and distance. Image-only or empty embeddings fall back to a neutral prior.
  5. Model capability catalog — The catalog defines candidate model capabilities, costs, limits, skill priors, and supported reasoning effort. Priors initialize expected success; they are assumptions to refine, not measured outcome guarantees.
  6. Eligibility filter — Hard capability and availability failures exclude a model. Soft restrictions and cost/latency limits may enter a relaxed pool only when a strict candidate pool is empty. The plan records exclusions and relaxation.
  7. Success curve — SuccessProb is a logistic item-response model. Effective model ability and effort are compared with request difficulty; this estimated probability is used by the objective and quality floor.
  8. Utility optimizer — Eligible model/effort arms are scored and ranked. choose prefers the highest-utility arm meeting the quality floor; if none meets it, it selects a low-cost arm close to the highest predicted probability.
  9. Session continuity — Route derives a session key and conversation fingerprints, then supplies any prior session state to Optimize. This enables prefix-aware costing and makes switching decisions sensitive to conversational continuity.
  10. Streaming dispatcher — Dispatcher builds an ordered arm list from the chosen candidate and fallbacks. It tracks attempted models, applies the retry budget, consults health state, and manages provider failures before commitment.
  11. Provider adapters — Native adapters translate the canonical request and response events into each provider protocol. The dispatch layer consumes their common Provider and Stream interfaces.
  12. First-token boundary — The Sink contract commits exactly once before sending the first event. Buffering upstream events until a visible response protects the client from partial output from a failed initial attempt.
  13. Provider health — Health tracking supplies model availability and observed speed to routing and dispatch. The decision environment incorporates this state before selecting a provider and attempt.
  14. Outcome learning — The learner consumes feedback signals to refine model ability estimates. When present, its Ability function is injected into the optimizer rather than leaving the catalog priors permanently fixed.
  15. Spend controller — The router reads the current budget controller multiplier and supplies it as Lambda to Optimize. This makes the cost penalty responsive to configured spending pressure.
03 / INTERACTIVE STUDY

Move the routing frontier

Tune task difficulty, the cost penalty, and the required quality estimate. Compare illustrative model arms on the same success curve used by the source.

CHANGE THE INPUTS

Illustrative abilities and costs, not provider benchmarks. Predicted success is not measured accuracy. The real optimizer may choose a best-available arm below the requested floor if no candidate meets it.

ILLUSTRATIVE MODELLIVE

04 / ENGINEERING CHOICES

Why it works this way.

01

Avoid an extra routing model call

Static embeddings and local predictors make the routing decision independent of a separate LLM round trip. Their limitation is that semantic similarity does not reliably distinguish an easy question from a deceptively hard one.

02

Make uncertainty and relaxation visible

The plan records difficulty, excluded candidates, objective values, and relaxed constraints. A quality floor is a preference over estimated probabilities, not a guarantee that a task will succeed.

03

Delay commitment to preserve options

A stream is buffered before the first visible token. Once response bytes are committed, changing the winning provider would violate the client’s coherent response stream.

05 / OPEN THE SOURCE

Trace it back.

Implementation details, examples, and project documentation.

Scope & limitations

  • The README’s success figures are simulations based on hand-set ability priors, not measured task completion on real traffic. This page makes no cost-saving or latency benchmark claim.
  • Provider catalogs and upstream capabilities change. Actual execution requires working credentials, compatible models, and the embedding assets; none are called by this portfolio lab.

Architecture and descriptions reflect the linked repository snapshot. The playground explains a mechanism; it does not execute the repository or report measured performance.