Interview problem: several product teams use different model providers. Create one governed API that enforces identity, budget, model eligibility, regional processing constraints and reliable usage accounting.
This is an illustrative design. Targets and prices are assumptions. A gateway can enforce an application policy; it cannot make every upstream model behave identically.
1. Scope and requirements
Clarify whether the gateway handles text only, tools, streaming, images, batches or embeddings. Assume text generation with optional structured output, two approved providers and one internal model. Tools execute in the caller's application, not in the gateway.
Functional requirements
- Authenticate applications and derive their tenant, use case and policy.
- Route requests only to eligible models and processing regions.
- Enforce request, token and spending limits before dispatch.
- Normalize common response fields while retaining provider-specific capabilities explicitly.
- Support streaming, cancellation, bounded retries and usage attribution.
- Version routes and policies, audit changes and roll back a rollout.
Non-functional requirements
- Add less than 50 ms p95 gateway overhead for ordinary admitted requests, excluding upstream generation.
- Target 99.95% monthly gateway availability; report end-to-end model availability separately.
- Prevent unauthorized fallback or cross-tenant cache reuse.
- Bound financial exposure during concurrent requests and uncertain provider outcomes.
- Keep raw prompts out of routine telemetry; redact and restrict sampled debugging records.
2. Estimate the workload
At 500 requests/s, mean 2,000 input and 400 output tokens, the gateway forwards 1M input tokens/s and eventually 200,000 output tokens/s at steady state. A ten-second mean request duration implies about 5,000 concurrent upstream requests. The stream relay must handle many open connections; CPU utilization alone is not the capacity signal.
Assume an eligible route's hypothetical price is $2/M input and $10/M output. A typical call costs .004 + .004 = $0.008; 500/s sustained for an hour is about $14,400. This is why budget enforcement cannot wait for an end-of-day dashboard. The rate is a peak assumption, not a forecast of monthly spend.
3. Baseline and its flaws
Read diagram source
flowchart LR
A[Product applications] --> G[Authenticated proxy]
G --> P[One model provider]
P --> G
G --> L[(Usage log)]
Start with a fixed allowed route and schema validation. The first experiment should establish pass-through overhead, error semantics and accounting accuracy.
| Failure | Repair | Benefit | Added cost |
|---|---|---|---|
| Products retry while proxy also retries | One attempt budget across layers | Contains amplification | Shared deadline/attempt contract |
| Provider outage triggers illegal region fallback | Filter eligibility before ranking | Preserves processing rules | Some requests must fail |
| Concurrent requests overspend one quota | Atomic reservation before dispatch | Bounded committed exposure | Reservation store and settlement |
| Streams end without usage metadata | Mark uncertain usage and reconcile | Honest accounting | Temporary budget holds |
| Common API hides unsupported features | Capability negotiation and validation | Predictable semantics | Less universal abstraction |
4. Detailed architecture
Read diagram source
flowchart TD
APP[Applications] --> AUTH[Identity and request validation]
AUTH --> POL[Policy and capability filter]
POL --> RES[Atomic budget reservation]
RES --> ROUTE[Deadline-aware route choice]
ROUTE --> A[Provider A adapter]
ROUTE --> B[Provider B adapter]
ROUTE --> C[Internal serving adapter]
A --> STREAM[Stream normalization and cancellation]
B --> STREAM
C --> STREAM
STREAM --> APP
A --> EVT[Append usage and attempt events]
B --> EVT
C --> EVT
EVT --> SETTLE[Idempotent settlement and reconciliation]
SETTLE --> LEDGER[(Budget ledger)]
LEDGER --> RES
ADMIN[Approved policy changes] --> CFG[(Versioned routing snapshot)]
CFG --> POL
CFG --> ROUTE
EVT --> OBS[Quality, latency, cost and error telemetry]
The gateway's data plane and control plane have different failure requirements. An operator dashboard outage should not stop all generation; revocation of a compromised application must still take effect within the agreed bound.
5. API and records
POST /generations includes an application-visible model alias, messages, supported response schema, output cap, idempotency key and deadline. The caller cannot choose an arbitrary upstream URL or set its own tenant identifier.
| Entity | Stored fields | Why |
|---|---|---|
| Policy snapshot | ID, tenant, model/region allowlist, expiry, limits | Reproduce an eligibility decision |
| Attempt | request ID, attempt number, provider, model revision, start/end, outcome | Separate logical work from paid calls |
| Reservation | tenant, request ID, maximum exposure, state, expiry | Prevent parallel overspend |
| Settlement | attempt ID, measured input/output, charged units, evidence | Deduplicate usage and reconcile bills |
A reservation must cover the permitted retry/candidate budget, not merely the cheapest first attempt. Where providers expose different tokenizers, estimate conservatively using the selected route and settle measured usage. A hard spend guarantee requires contracts for unknown or delayed charges; state the bounded uncertainty instead of claiming exact real-time billing.
6. Route and execute
- Authenticate and determine permitted use-case policy.
- Reject unsupported modalities, schemas, token limits or destinations.
- Reserve the maximum allowed attempt budget atomically against the tenant allowance.
- Rank only eligible routes by a tested quality/latency/cost policy.
- Pin the route snapshot and dispatch with the original deadline.
- Relay events with backpressure. Retain the actual model identity and finish reason.
- Settle known usage once, hold uncertain exposure and reconcile provider records.
A fallback is a new model behavior. Test its quality, tool-call shape and refusal/error behavior. Do not silently label its answer as the originally requested model's output.
7. Retries and degradation
| Situation | Behavior |
|---|---|
| Rejected before dispatch | Release reservation; no inference retry needed |
| Rate limited before any output | Retry eligible route only within deadline and reserved budget; use jitter |
| Timeout with unknown completion | Record uncertainty; avoid assuming zero charge |
| Partial stream delivered | Explicit interruption or retained-event resume; no invisible replacement answer |
| Budget store unavailable | Fail closed for spend-controlled traffic or use a separately approved bounded allowance |
| Provider quality regresses | Disable route by evaluated signal; preserve a known approved version where available |
A circuit breaker limits repeated failing calls. It is not a quality detector; monitor product-specific evaluation as well as HTTP status.
8. Cost and benefit
For an illustrative one million successful tasks per month, lowering average complete-task cost from $0.014 to $0.010 saves $4,000 before gateway operating cost. If the gateway costs $3,000 monthly and adds $2,000 of allocated maintenance, the change loses $1,000 on that accounting boundary. Governance or reliability may still justify it; call that benefit explicitly.
| Decision | Benefit | Tradeoff |
|---|---|---|
| Static routing first | Easy to audit and reproduce | Less adaptive optimization |
| Learned routing later | Can match difficulty to cost | Router evaluation, drift and failure modes |
| Exact response cache | Low-cost repeated safe requests | Invalidation and identity/version keys |
| Semantic cache | More reuse | False-equivalence risk; needs separate evaluation |
| Multi-region gateway | Better local resilience | Policy propagation and budget consistency |
9. Release and interview follow-ups
Test concurrent quota depletion, duplicate settlement, provider timeouts, a revocation during a stream, unavailable policy storage and an unsupported schema. Canary one application; compare complete-task quality, total cost and p95 overhead before expanding.
Q1: Why can the cheapest model be an expensive route?
Sample answer: It may need more retries, longer outputs, human repair or escalation. I compare cost per acceptable task, including routing and fallback calls, against a fixed quality and latency requirement.
Q2: Can provider compatibility remove vendor differences?
Sample answer: It can normalize a useful subset. It cannot guarantee identical tokenization, tool-call semantics, structured-output support, safety behavior, context limits or billing. The contract should expose unsupported capabilities explicitly.
Q3: What must an outage fallback preserve?
Sample answer: Authorization, regional/data-processing restrictions, required capabilities, budget and a tested minimum quality. If no route satisfies those constraints, fail explicitly instead of sending data to an unapproved service.
Closing remarks
I would ship a thin governed gateway with static approved routes, atomic reservations, explicit streaming semantics and reconciled usage. Adaptive routing comes after a reliable evaluation loop. The key tradeoff is centralized control versus an additional critical dependency.
| Recall | Explain |
|---|---|
| Filter before rank | Policy determines eligibility |
| Reserve before send | Concurrency can overspend |
| Count attempts and tasks | A retry is still work and possibly cost |
| Fail explicitly | An ineligible fallback is not resilience |
Tip: Separate gateway overhead from upstream generation latency. Combining them hides which system must improve.