Learnastra AI SYSTEM DESIGNAnup Rai

Concept · Understand the mechanism

AI gateways and model routing: enforce policy before choosing a model

By Anup Rai8 min readReviewed September 2026

An AI gateway is an intermediary on the request path between applications and model services. It can centralize authentication, policy, routing, usage accounting and traffic controls. Its data plane handles requests and responses; its control plane manages configuration such as provider credentials, routes and budgets. Calling the whole gateway only a control plane misses the request-serving responsibility.

Model routing selects a model or endpoint for a task. Load balancing distributes work among eligible endpoints. Fallback selects another path after a failure. These mechanisms can work together, but they solve different problems.

When the extra layer earns its cost

A small application can begin with a direct SDK behind a narrow application interface. A shared gateway becomes useful when several applications duplicate access, accounting, quota or routing logic, or when centralized policy is a concrete requirement.

The number of providers is not a universal adoption threshold. One provider serving many teams may justify a gateway. Two providers serving a small prototype may not. The added component needs availability, capacity and an operating owner.

Responsibility Gateway can centralize Still needs application context
Authentication Verify an application/virtual key End-user identity and resource authorization
Routing Enforce eligible models/endpoints Required task capabilities and quality
Traffic Rate/concurrency limits, deadlines, retries Whether delayed or degraded output is acceptable
Accounting Attempts, usage and rate references What counts as a successful user outcome
Caching Store/reuse eligible responses Access, freshness, policy and semantic validity
Filtering Configured input/output controls Domain rules and authorization at actual tool execution

A common API shape reduces adapter work; it does not guarantee identical tool calls, structured output, tokenization, streaming, reasoning or error semantics across providers.

Interview design: shared access for the learning platform

Assume lesson feedback, quiz explanation and offline content checks share model access.

Functional requirements

  1. Authenticate each application and propagate trusted account/workload scope.
  2. Select only endpoints approved for the request's capabilities and data policy.
  3. Apply request, token, concurrency and spend controls.
  4. Stream or return results while recording all attempts and their outcomes.
  5. Support bounded fallback and versioned routing changes.

Non-functional requirements

  1. Prevent clients from overriding mandatory provider/access restrictions.
  2. Bound gateway overhead and preserve end-to-end deadlines.
  3. Avoid retry storms and cross-account cache leakage.
  4. Keep accounting and request correlation durable enough for reconciliation.
  5. Remain available through a gateway-replica failure and define behavior when shared policy services fail.

Start with explicit routes: quiz explanations and feedback use evaluated configurations; offline jobs use a separate quota class. Add learned routing only if its quality/cost benefit exceeds the extra inference, evaluation and operating complexity.

Architecture / visual model
flowchart TD A[Application and trusted request scope] --> G[Gateway replicas] P[Versioned policy and provider registry] --> G G --> E[Eligibility: access, data, capabilities] E --> B[Reserve quota, concurrency and budget] B --> C{Eligible cached result?} C -->|Yes| O[Return with provenance] C -->|No| R[Choose eligible endpoint] R --> M[Provider or self-hosted model] M --> H{Outcome} H -->|Success| O H -->|Retryable and budget remains| R H -->|Final or uncertain| F[Explicit failure or recovery path] O --> U[Reconcile usage and release reservations] F --> U
Read diagram source
flowchart TD
    A[Application and trusted request scope] --> G[Gateway replicas]
    P[Versioned policy and provider registry] --> G
    G --> E[Eligibility: access, data, capabilities]
    E --> B[Reserve quota, concurrency and budget]
    B --> C{Eligible cached result?}
    C -->|Yes| O[Return with provenance]
    C -->|No| R[Choose eligible endpoint]
    R --> M[Provider or self-hosted model]
    M --> H{Outcome}
    H -->|Success| O
    H -->|Retryable and budget remains| R
    H -->|Final or uncertain| F[Explicit failure or recovery path]
    O --> U[Reconcile usage and release reservations]
    F --> U

The actual ordering of cache lookup and reservation depends on which work is billed and rate-limited. Authentication, access and cache eligibility must happen before a cached answer is exposed. Every fallback repeats the required eligibility checks.

Eligibility first, optimization second

A route decision should first exclude endpoints that violate a hard requirement:

  1. Required model modality, context capacity, tool or output-schema support.
  2. Permitted data processing location and provider/data-retention policy.
  3. Account access and organizational model restrictions.
  4. Available deadline, quota and budget.
  5. Required evaluated task quality.

Then rank eligible choices by the product's objective. A cheap endpoint outside the permitted region is not a valid candidate. If none qualify, return a clear unavailable result or an approved degraded mode; do not silently relax constraints.

Routing approach Useful when Main tradeoff
Static/task rule Task classes and requirements are known Simple and explainable; rules need maintenance
Cost-aware Several choices meet quality/capability requirements Must include retries, evaluation and output-length effects
Latency/load-aware Endpoint conditions vary Estimates can be stale; avoid sending everyone to the same endpoint
Semantic/learned router Task difficulty or type is hard to describe with rules Router inference, errors and distribution drift
LLM classifier Classification needs model judgment Adds a model call, cost and another failure path
Cascade A first result can be checked and escalated Sequential latency and payment for both attempts

RouteLLM studies learned routing between stronger and weaker models. Its reported results belong to particular models and evaluations. Routing before generation differs from a cascade that generates a cheaper answer and then decides whether to escalate. Neither inherently supplies outage recovery.

Calculate cascade economics

With hypothetical costs of $0.002 for a first model, $0.020 for an escalated model, and $0.001 for a verifier per task:

expected cost = first call + verification + escalation rate × second call
at 20% escalation: 0.002 + 0.001 + 0.20 × 0.020 = $0.007
at 90% escalation: 0.002 + 0.001 + 0.90 × 0.020 = $0.021

The second case costs more than using the stronger model once, before adding routing overhead. Compare accepted quality and latency as well as dollars. Model-reported confidence is not automatically calibrated; a weak verifier can accept precisely the answers that needed escalation.

Retry errors by meaning, not only status family

Outcome Typical response
Temporary capacity/rate limit Respect provider guidance, back off or use an independently eligible endpoint
Transient service/network error Bounded retry if the request is safe to repeat and time remains
Invalid schema/context too long Correct or reject the request; repeated identical calls usually do not help
Authentication/authorization failure Fix configuration or deny access; do not bypass the restriction
Unknown model/deployment Treat as configuration/lifecycle issue unless a tested mapping explicitly handles it
Safety/policy refusal Follow application policy; do not route around the restriction merely to obtain an answer
Timeout after a possible external effect Reconcile the effect before repeating it

429 is itself a 4xx status, so “never retry any 4xx” is too broad. Conversely, not every 5xx warrants another attempt beyond the deadline. Preserve meaningful error details without exposing credentials or private content.

Use an end-to-end attempt/time budget, exponential backoff with jitter where appropriate, and Retry-After guidance when supplied. Avoid multiplying retries at every layer. One application retry around three gateway attempts can already make six provider attempts.

A circuit breaker stops sending normal traffic to a failing dependency while it recovers, then probes cautiously. Scope health to the relevant endpoint/account/limit: one tenant's invalid key or quota exhaustion need not disable service for everyone.

More keys do not automatically create more quota. Limits may be shared at account, organization, region or deployment level. Plan legitimate capacity; do not treat key rotation as a way to bypass provider limits.

Streaming changes fallback behavior

Before output reaches the client, a safe generation-only request may be retried under policy. After a partial answer is emitted, appending a different model's fresh answer can create contradictory text or malformed tool/JSON output. End the stream with an explicit failure, or use a designed restart/resume protocol that tells the client what to replace.

A tool action may have succeeded even if its response was lost. A new model or provider does not make that action safe to repeat. Use operation IDs and effect reconciliation.

Current product options and what to inspect

Documentation reviewed in September 2026:

Option Relevant surface Adoption check
LiteLLM Router/proxy controls and provider adapters Retry ownership, caller overrides, supported error semantics and shared-state operations
OpenRouter Managed provider selection and restrictions Required-parameter support and actual provider/data-policy eligibility
Cloudflare AI Gateway Managed gateway, caching, rate controls and routing features Feature maturity, including beta spend-limit/dynamic-routing surfaces
Portkey Gateway routing and operational controls Hosting mode, policy scope, data paths and contractual limits
Kong AI Gateway AI capabilities within an API gateway platform Required plugins, editions, supported providers and operating fit
Agent Router Current destination of Envoy AI Gateway's provider-fallback documentation Supported routing/retry behavior and deployment integration

For example, OpenRouter documents require_parameters for excluding providers that do not support all requested parameters. Do not assume every adapter refuses unsupported options by default. LiteLLM documents separate retry configuration and caller overrides; enforce product limits at a boundary the caller cannot weaken. Verify the installed version rather than copying an old comparison table.

Managed options reduce infrastructure work but still require policy, cost and incident ownership. Self-hosting the proxy does not keep model payloads local when the next hop is an external provider.

Availability, budgets and observability

Run appropriate redundant gateway replicas and avoid keeping authoritative budgets only in process memory. Reserve estimated spend atomically before concurrent work; reconcile all attempts afterward. Decide whether a missing budget/policy dependency means fail closed, a restricted preallocated allowance, or another explicit behavior.

An alerting dashboard is not a hard spend cap. Provider usage may arrive late, and cancellation does not guarantee zero further charges. See FinOps controls.

Record request/task ID, account/workload scope, route-policy revision, attempted endpoints, outcome, latency, token usage and estimated/reconciled cost. Redact before exporting content. Use bounded dimensions for aggregate metrics; keep per-request identifiers in traces.

Observed problem Repair Cost/benefit
Every retry hits the same exhausted quota Model quota domains and cool down the affected scope More accurate availability state
Alternate provider violates required behavior Capability contracts and paired evaluations Fewer eligible fallbacks
Cache hit leaks private context Scope by access, relevant versions and freshness Reduced hit rate but correct results
Gateway is a new single point of failure Redundant data plane and tested shared dependencies Extra operating cost
Latency router creates a traffic stampede Load-aware selection and controlled exploration More routing complexity
Retried jobs exceed the tenant budget Shared reservation and total-attempt accounting Persistent state and reconciliation

Interview questions and answer checks

  1. Is a gateway a control plane? It usually includes control-plane configuration and a data plane that serves requests.
  2. Can model routing be useful with one provider? Yes, if different models or deployments serve different requirements; centralized access/accounting may also help.
  3. Why can a fallback hurt reliability? It may lack capacity, change behavior, violate policy or repeat an uncertain effect.
  4. Do multiple API keys multiply quota? Not necessarily; inspect the provider's quota scope.
  5. Why not retry every error? Permanent/configuration failures waste capacity, and uncertain writes can be duplicated.
  6. What must happen before semantic cache reuse? Check access, task suitability, freshness and relevant versions; similarity alone is insufficient.
  7. How do you validate a learned router? Compare end-to-end quality, cost, latency and important slices against a simple baseline; monitor distribution drift.
  8. When is a gateway overkill? When the application can implement the required boundaries clearly with less operational complexity; provider count alone is not decisive.

Final notes

Remember eligible → selected → attempted → accounted. Preserve policy through every route and retry. Centralize controls when doing so makes the system easier to operate, and prove that the new shared layer can meet the reliability requirements it inherits.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← CI/CD for LLM applications: release the complete behavior
NEXT LESSONFinOps and token economics: measure cost per useful outcome →

Explore the diagram