Predict locally
Structural request features combine with a pure-Go Model2Vec embedder. A nearest-neighbor exemplar bank and ridge head estimate difficulty, skill mix, output length, and uncertainty without calling a routing LLM.
An OpenAI-compatible gateway that chooses a model and reasoning effort for each request. A local predictor estimates task difficulty and skill requirements; a constrained optimizer balances expected success, cost, and latency before a streaming dispatcher handles provider execution and failover.
A model name is a coarse allocation decision. ModelRouter treats model and effort together as an arm, estimates how well that arm fits the request, and makes the cost of a stronger answer explicit while preserving protocol compatibility for clients.
Structural request features combine with a pure-Go Model2Vec embedder. A nearest-neighbor exemplar bank and ridge head estimate difficulty, skill mix, output length, and uncertainty without calling a routing LLM.
The optimizer filters model candidates by capabilities, credentials, and health, then evaluates each supported effort. Utility combines predicted success with cost and latency penalties; session state influences switching and cache-aware cost.
Upstream responses are streamed into a peek buffer. The dispatcher can switch attempts before the first visible token commits a stream, with effort-scaled timeouts, retry budgets, and optional hedging.
Feedback and observed response signals update per-skill model offsets and exemplars. Spend control adjusts the cost penalty, while health tracking informs later routing decisions.
Local selection is separate from upstream execution; feedback reconnects the two.
Tune task difficulty, the cost penalty, and the required quality estimate. Compare illustrative model arms on the same success curve used by the source.
Illustrative abilities and costs, not provider benchmarks. Predicted success is not measured accuracy. The real optimizer may choose a best-available arm below the requested floor if no candidate meets it.
Static embeddings and local predictors make the routing decision independent of a separate LLM round trip. Their limitation is that semantic similarity does not reliably distinguish an easy question from a deceptively hard one.
The plan records difficulty, excluded candidates, objective values, and relaxed constraints. A quality floor is a preference over estimated probabilities, not a guarantee that a task will succeed.
A stream is buffered before the first visible token. Once response bytes are committed, changing the winning provider would violate the client’s coherent response stream.
Implementation details, examples, and project documentation.
Nearest-neighbor/ridge blending, uncertainty, and logistic success estimation.
Hard and soft filters, effort candidates, quality-floor handling, and fallbacks.
Retry budgets, attempt state, stream commitment, timeouts, and provider failover.
Explicitly distinguishes simulated routing success from real-traffic measurements.
Architecture and descriptions reflect the linked repository snapshot. The playground explains a mechanism; it does not execute the repository or report measured performance.