Pricing is the schedule of charges for a service; cost is the charge incurred by a measured workload. Total cost of ownership (TCO) includes the infrastructure, people and operating work within an explicitly stated accounting boundary. An API rate is one input to that calculation.
Rate-card review: September 24, 2026. Unless another unit is shown, table prices are USD per million tokens (MTok), using direct-provider Standard list rates. They exclude taxes, negotiated discounts, marketplace terms and regional/service modifiers unless stated. Current prices and announced future prices are separate. Workload sizes, staffing and comparison costs are interview assumptions.
| Remember the unit | Example | Common mistake |
|---|---|---|
| Rate × quantity | USD 2/MTok × 0.002 MTok = USD 0.004 | Treating 2,000 tokens as two million |
| Call vs task | Five model calls can serve one customer task | Multiplying customer tasks by one call’s bill |
| Usage vs outcome | Failed attempts can still consume billed resources | Dividing only successful-attempt spend by successes |
| Estimate vs invoice | Logged usage is rated, then reconciled | Treating an incomplete stream as zero cost |
| Total vs marginal cost | Idle reserved GPUs still cost money | Adding paid idle capacity a second time |
Read tokenization, capability assessment and model selection alongside the calculations.
Token units and rates
A token is a unit produced by a model's tokenizer, not a fixed number of words. Two models can tokenize the same text differently. A price quoted per million tokens must be multiplied by the measured token count divided by one million. Input and output can have different rates, and the provider's usage fields determine which categories are billable.
Use invented round numbers to practice. At $2 per million input tokens, 2,000 input tokens cost 2,000 / 1,000,000 × $2 = $0.004. At $10 per million output tokens, 500 output tokens cost $0.005. The call costs $0.009 before tools, storage, or other charges. One hundred thousand such calls cost $900. That is a call forecast; if each user task makes five calls, the task volume is different.
Partition cached usage without double-counting
Suppose a provider reports 10,000 total input tokens, of which 8,000 are eligible cache reads. In an illustrative contract, uncached input costs $2 per million, cache reads cost $0.20, and 1,000 output tokens cost $10 per million. The input is 2,000 uncached tokens plus 8,000 cached tokens, not 10,000 uncached plus another 8,000 cached. The bill is $0.004 + $0.0016 + $0.01 = $0.0156 before any separate cache creation or storage charge.
Some providers report usage categories differently, so map the actual fields to non-overlapping billed quantities. Cache writes may be a separate category or priced under a documented multiplier. Reasoning usage may already be included in output billing. Adding it again overstates cost. The exact contract belongs beside the calculation, not in an assumption hidden inside a spreadsheet.
Provider prompt caching reuses eligible prompt processing. An answer cache returns a previous answer. The latter requires an additional correctness decision: do the query, user permissions, source versions, and freshness make reuse appropriate? Similar wording alone is not enough.
From a rate card to a workload forecast
Multiply by the distribution of tasks, not one unusually short example. Include tool calls, retries, escalations, long-context tiers, modalities, and service options that apply. Separate assumptions about present prices from announced future changes and from rates that could not be verified. A promotional price is not a permanent architectural guarantee.
Then add non-model costs: ingestion, embeddings, indexes, storage, network, observability, human review, and operating capacity. Define which costs belong to the business comparison. A self-hosted model's raw accelerator rental does not include all of those responsibilities, and an API's token bill does not include all application operations either.
The rate tables below are reference material. For an interview, memorize the arithmetic and the cost drivers. Explain which provider facts you would verify before making a purchasing or design decision.
Table of Contents
- What Appears on the Bill
- Current Text API Pricing
- Specialist APIs and Tool Fees
- Cost Calculation
- Worked Examples
- Context Caching Economics
- Cost Optimization
- Self-Hosting Economics
- Total Cost of Ownership and FinOps
- API versus self-hosted worksheet
- Forecast boundary cases
- Full FinOps design interview
- Interview questions
- Official Sources
What Appears on the Bill
A production request can create several independently billed quantities:
| Cost category | What to measure |
|---|---|
| Uncached input | System instructions, user input, retrieved context, images converted to tokens, tool schemas, and prior conversation state |
| Cached input | Reused prompt-prefix tokens actually reported as cache hits |
| Cache writes and storage | Tokens written into a cache, the cache duration, and any token-hour storage fee |
| Output and reasoning | Visible output plus any provider-billed reasoning or thinking tokens |
| Tools | Search queries, retrieval calls, code containers, computer use, and other server-side tools |
| Media | Audio minutes or tokens, generated images, video seconds, OCR pages, and transcription |
| Service and region | Batch, flex, fast/priority, data-residency, or marketplace modifiers |
| Failure overhead | Retries, fallbacks, agent loops, shadow traffic, evaluation traffic, and partial failures |
The general monthly model is:
where token quantities are measured in millions and the non-token fees are monthly totals counted once. Media converted to billed model tokens belongs in the token term; use separate media fees only for distinct charges. Add serving infrastructure, observability, human review, support, and incident costs when comparing total cost of ownership.
There is no universal enterprise discount schedule. Volume discounts, committed capacity, support, regional routing, and cloud-marketplace terms are provider- and contract-specific. Use a written quote for a procurement model.
Current Text API Pricing
OpenAI text models
| Model | Input | Cached input | Cache write | Output |
|---|---|---|---|---|
gpt-6-astra |
$10.00 | $1.00 | $12.50 | $50.00 |
gpt-6-sol |
$2.00 | $0.20 | $2.50 | $10.00 |
gpt-6-luna |
$0.10 | $0.01 | $0.125 | $0.50 |
gpt-5.6-sol — older model |
$4.00 | $0.40 | $5.00 | $20.00 |
gpt-5.6-terra — older model |
$2.00 | $0.20 | $2.50 | $12.00 |
gpt-5.6-luna — older model |
$0.20 | $0.02 | $0.25 | $1.20 |
For these models, prompts above 272K input tokens use whole-request long-context rates: input, cache read and cache write are 2× the short-context rates; output is 1.5×. For example, GPT-6 Sol becomes $4 / $0.40 / $5 / $15 in the table’s column order. This is not a surcharge on only the excess tokens.
Batch and Flex use half the listed token rates; Fast uses twice Standard. Check availability as well as price: GPT-6 Astra, Sol and Luna support EU data residency only with Standard processing. Eligible regional processing adds 10%. GPT-5.6 Sol’s older promotional rate is guaranteed at least through November 21, 2026; that is not an announced future increase.
Astra requires reasoning; billed reasoning can exceed the visible answer. Its short-context input/output rates are 5× GPT-6 Sol’s, but different token consumption and retry behavior determine the actual task cost. OpenAI rate card.
Anthropic Claude
| Model | Base input | Cache read | 5-minute write | 1-hour write | Output |
|---|---|---|---|---|---|
| Claude Fable 5.1 | $10.00 | $0.25 | $12.50 | $20.00 | $50.00 |
| Claude Opus 5.5 | $4.00 | $0.20 | $5.00 | $8.00 | $20.00 |
| Claude Sonnet 5 | $2.00 | $0.20 | $2.50 | $4.00 | $10.00 |
| Claude Haiku 4.5 | $1.00 | $0.10 | $1.25 | $2.00 | $5.00 |
Cache reads are 2.5% of base input for Fable 5.1, 5% for Opus 5.5, and 10% for Sonnet 5/Haiku 4.5. Restricted-access Mythos 5.1 shares Fable’s rates. Do not apply one cache multiplier to every Claude model.
Batch halves input/output rates. Claude 4.6-and-later long context has no higher token tier. Opus 5.5 Fast costs $8 input / $40 output, is first-party-only and does not combine with Batch. US-only inference_geo on eligible models adds 10% to token categories. Newer tokenizers can change counts for the same text. Claude pricing.
Google Gemini
| Model | Standard input | Cached input | Cache storage | Standard output |
|---|---|---|---|---|
gemini-3.8-flash / gemini-3.7-flash through Dec. 31, 2026 |
$0.75 | $0.075 | $0.50 / MTok-hour | $3.75 |
gemini-3.8-flash / gemini-3.7-flash starting Jan. 1, 2027 |
$1.50 | $0.15 | $1.00 / MTok-hour | $7.50 |
gemini-3.5-flash |
$1.50 | $0.15 | $1.00 / MTok-hour | $9.00 |
gemini-3.5-flash-lite |
$0.30 | $0.03 | $1.00 / MTok-hour | $2.50 |
Google's output price includes thinking tokens. Batch and Flex input/output rates are half Standard for the models above; check their separate cache and storage columns rather than applying one blanket multiplier. Gemini 3.8 is the latest stable Flash; 3.7 remains stable at the same listed price. Paid-tier content is not used to improve Google's products; the free tier has different data-use terms.
gemini-3.1-pro-preview has a two-tier standard rate: $2 input / $12 output for
prompts up to 200K tokens, and $4 / $18 above 200K. It is a preview model, so
its lifecycle risk belongs in the decision alongside price.
xAI, Mistral, and DeepSeek
| Provider | Model | Input or cache miss | Cached input | Output |
|---|---|---|---|---|
| xAI | grok-4.7 |
$2.00 | $0.50 | $6.00 |
| Mistral | Mistral Large 3 | $0.50 | $0.05 | $1.50 |
| Mistral | Mistral Medium 3.5 | $1.50 | $0.15 | $7.50 |
| Mistral | Mistral Small 4 | $0.15 | $0.015 | $0.60 |
| Mistral | Ministral 3 14B | $0.20 | $0.02 | $0.20 |
| Mistral | Ministral 3 8B | $0.15 | $0.015 | $0.15 |
| Mistral | Ministral 3 3B | $0.10 | $0.01 | $0.10 |
Grok 4.7 has a 500K context window. At 200K prompt tokens or more, its whole-request rates become $4 input / $1 cached input / $12 output. xAI's current rate card now publishes both tiers explicitly. The US regional endpoint adds 10% to token rates.
DeepSeek has separate peak and off-peak rates:
| API model / current version | Peak input miss | Peak cache hit | Peak output |
|---|---|---|---|
deepseek-flash / V4.1 Flash |
$0.30 | $0.006 | $1.20 |
deepseek-v4-pro / V4 Pro 0813 |
$1.32 | $0.044 | $3.96 |
Off-peak is half of every listed rate. Peak windows are Monday–Friday 01:00–04:00 and 06:00–10:00 UTC, excluding Chinese public holidays. Legacy Flash names route to V4.1 Flash; V4 Pro remains separately listed. Record the actual version, billing window and usage. DeepSeek rate card.
Mistral figures above use its global Standard schedule; use the actual regional and service contract when budgeting. Mistral pricing.
Why the cheapest token row may lose
Rate tables do not normalize:
- provider tokenizers or media-to-token conversion;
- reasoning-token consumption;
- tool-schema and system-prompt overhead;
- successful-task rate, retries, or human correction;
- latency, capacity, and service tier;
- safety, governance, support, or data placement.
The decision metric should usually be cost per successful task at the required quality and service level, not cost per token.
Specialist APIs and Tool Fees
Do not compare unlike units as if they were all token prices.
| Provider / endpoint | Billing unit | Current standard list price |
|---|---|---|
OpenAI gpt-live-1 |
Session minute, billed per second | $0.05; backend models/tools extra |
OpenAI gpt-realtime-2.1 audio |
MTok input / cached / output | $32 / $0.40 / $64 |
OpenAI gpt-realtime-2.1-mini audio |
MTok input / cached / output | $10 / $0.30 / $20 |
OpenAI gpt-transcribe |
Estimated audio minute | $0.0045 |
OpenAI gpt-live-transcribe |
Estimated audio minute | $0.017 |
OpenAI gpt-realtime-translate |
Estimated audio minute | $0.034 |
OpenAI gpt-image-2.5-sunburst / gpt-image-2.5-flare image tokens |
MTok input / cached / output | $8 / $2 / $30 |
OpenAI text-embedding-3-small / text-embedding-3-large |
MTok input | $0.02 / $0.13 |
| Mistral Codestral Embed | MTok input / cached input | $0.15 / $0.015 |
Google gemini-3.5-transcribe |
Estimated blended minute | approximately $0.005 |
xAI grok-voice-think-fast-2.0 |
Audio minute / text input | $0.08 / $0.004 |
| xAI speech-to-text | Audio hour | $0.10 REST; $0.20 streaming |
xAI grok-imagine-image-2.0 |
Output image, 1K–2K low/medium | $0.04–$0.08; image input adds $0.01/image |
xAI grok-imagine-video-1.5 |
Generated second, 480p / 720p / 1080p | $0.08 / $0.14 / $0.25; image input adds $0.01/image |
| Mistral OCR 4.1 | 1,000 pages | $4 standard; $0.40 cached |
| Mistral Voxtral Mini Transcribe 2 | Minute | $0.003 standard; $0.0003 cached |
GPT-Image-2.5 text input costs $5/MTok ($1.25 cached), separate from image tokens. Do not assume every image model supports Batch because an earlier model did. Use its explicit rate card and output settings to estimate each image.
The Videos API and Sora 2 models have a September 24, 2026 shutdown date and are excluded from new-workload budgets. A historical rate does not establish availability. OpenAI’s self-serve fine-tuning platform is also winding down and closed to new users; do not assume a current text model supports new fine-tuning jobs. OpenAI deprecations.
Embedding bills cover embedding computation, not the vector database, index rebuilding or document ingestion. Training/adaptation budgets additionally need data preparation, label review, training runs, evaluation and serving; see fine-tuning.
Common tool charges can be large enough to change model routing decisions:
| OpenAI tool | Current price |
|---|---|
| Web search | $10 / 1,000 calls, plus search-content tokens at the model rate |
| File search calls | $2.50 / 1,000 calls |
| File search storage | $0.10 / GB-day after 1 GB free |
| Hosted shell / code container | $0.03 for 1 GB, $0.12 for 4 GB, $0.48 for 16 GB, or $1.92 for 64 GB per 20-minute session |
OpenAI notes that eligible containers are billed per minute with a five-minute minimum, not necessarily one full 20-minute unit. Check actual session usage.
X Search is now billed by fetched items: $5/1,000 posts and $10/1,000 profiles. Parent and quoted posts count. It is no longer a flat per-call charge. xAI tool pricing.
Google Search grounding provides 5,000 free search requests per month shared across Gemini 3.x models on the paid tier, then charges $14 per 1,000 search queries. One customer request may trigger multiple queries, so log the billed query count. Other providers also charge for some server-side tools; verify the specific tool version and deployment path. Google pricing.
Cost Calculation
Use the provider's usage fields for metering and reconcile them with the invoice. Normalize provider-specific fields into non-overlapping billed categories; some APIs report cached input as a subset of total input. Keep rates in a dated configuration instead of hard-coding them throughout the application.
from dataclasses import dataclass
from decimal import Decimal
@dataclass(frozen=True)
class TokenRates:
input_per_mtok: Decimal
cached_input_per_mtok: Decimal
cache_write_per_mtok: Decimal
output_per_mtok: Decimal
def __post_init__(self):
for value in vars(self).values():
if not isinstance(value, Decimal) or not value.is_finite() or value < 0:
raise ValueError("rates must be finite non-negative Decimals")
def token_cost(rates, *, uncached_input=0, cached_input=0,
cache_write=0, output=0, extra_fees=Decimal("0")):
if not isinstance(rates, TokenRates):
raise ValueError("a validated rate card is required")
counts = (uncached_input, cached_input, cache_write, output)
if any(type(value) is not int or value < 0 for value in counts):
raise ValueError("token counts must be non-negative integers")
if (not isinstance(extra_fees, Decimal)
or not extra_fees.is_finite() or extra_fees < 0):
raise ValueError("fees must be a finite non-negative Decimal")
return (
uncached_input * rates.input_per_mtok
+ cached_input * rates.cached_input_per_mtok
+ cache_write * rates.cache_write_per_mtok
+ output * rates.output_per_mtok
) / Decimal("1000000") + extra_fees
GPT_6_SOL = TokenRates(*(Decimal(x) for x in ("2", "0.20", "2.50", "10")))
cost = token_cost(GPT_6_SOL, uncached_input=2_600, output=300)
print(f"USD {cost:.6f}") # USD 0.008200
This function covers token categories and explicitly supplied fees. It does not infer long-context tiers, service modifiers, cache-storage duration, media conversion, taxes, or contract terms. Select the correct rate card before calling it. Create decimals from strings, accumulate sub-cent usage, and round at the documented billing boundary. A production meter also needs bounded inputs, currency, immutable rate revisions and reconciliation records.
For an agent, calculate every turn and tool invocation. In this formula, model-call cost excludes the separately summed tool fees:
Also record abandoned runs. Excluding failures makes the cost per successful task look artificially low.
Worked Examples
RAG chatbot
Assume each request has 2,600 uncached input tokens and 300 total billed output tokens, including any billed reasoning tokens. There are no cache, tool, retry, or service-modifier charges. Actual reasoning runs may use much more output.
| Model | Cost per request | 10,000 requests/day | 30-day month |
|---|---|---|---|
| GPT-6 Astra | $0.041000 | $410.00 | $12,300.00 |
| GPT-6 Sol | $0.008200 | $82.00 | $2,460.00 |
| Claude Opus 5.5 | $0.016400 | $164.00 | $4,920.00 |
| Claude Sonnet 5 | $0.008200 | $82.00 | $2,460.00 |
| Gemini 3.8 Flash, 2026 price | $0.003075 | $30.75 | $922.50 |
| GPT-6 Luna | $0.000410 | $4.10 | $123.00 |
This is arithmetic, not a model recommendation. If a lower-rate model needs more retries, produces more output, or fails the quality threshold, its cost per accepted answer can be higher.
Document summarization
For 8,000 uncached input tokens and a 500-token output using GPT-6 Sol:
That is $21 for 1,000 documents or $210 for 10,000 documents, before storage, retrieval, retries, and review.
Tool-using agent
Assume one run makes eight GPT-6 Sol calls totaling 20,000 uncached input tokens, 25,000 cached input tokens, and 4,000 output tokens, plus five OpenAI web-search calls:
| Component | Cost |
|---|---|
| Uncached input | $0.040 |
| Cached input | $0.005 |
| Output | $0.040 |
| Five searches | $0.050 |
| Total | $0.135 per run |
At 1,000 runs per day, the illustration is $4,050 for a 30-day month. It excludes the initial cache write, search-content tokens, containers, failed runs, and any regional or service-tier modifier.
Context Caching Economics
Caching is valuable when a sufficiently large prefix repeats within the cache lifetime. It is not automatically valuable merely because a prompt is long.
For a reusable prefix of one MTok, let:
- be the ordinary input price;
- be the cache-write price;
- be the cache-read price;
- be total uses, including the initial write.
Ignoring storage fees, caching wins when:
when $B > R$:
For GPT-6/GPT-5.6 and Sonnet/Haiku 5-minute caches, and . Fable/Mythos 5.1 use ; Opus 5.5 uses . In these cases, the 1.25× write pays back on the second total use: one write plus one read. Claude's 1-hour cache uses and pays back on the third total use. These conclusions assume full-prefix hits, no storage fee, and reuse within the actual provider TTL; OpenAI's cache duration is not Claude's five minutes.
Google's explicit caching also charges token-hours of storage, so add:
Measure actual cache-hit tokens, reuse count, TTL expiry, prefix churn, and the cost of cache misses. Put stable instructions and shared documents before request-specific data when the provider's caching rules are prefix-based.
Cost Optimization
Optimize only after measuring quality and usage. The highest-leverage controls usually are:
- Measure cost per successful task. Join usage, tool fees, retries, human review, latency, and outcome quality by request or agent run.
- Shorten what is repeatedly sent. Remove redundant instructions, trim tool descriptions, summarize state, and retrieve only relevant context.
- Cache stable prefixes. Validate the realized hit rate and include cache writes and storage in the calculation.
- Use Batch or Flex for deferrable work. OpenAI, Anthropic, and Google publish 50%-lower token rates for eligible asynchronous/deferred paths, but their latency and feature contracts differ.
- Route with evaluations. Send easy work to a lower-cost model only after measuring routing errors, escalation frequency, and cost per passing result.
- Bound agents. Cap model turns, output, reasoning effort, tool calls, elapsed time, and dollars per run. Detect repeated state and tool loops.
- Use specialist endpoints. OCR, transcription, moderation, embeddings, and media generation often have better pricing and behavior than a general model forced into the task.
- Control failure amplification. Use bounded retries with jitter, idempotent tools, circuit breakers, and failure-aware fallbacks.
- Clean up billed storage. Expire unused vector stores, files, caches, and container sessions according to retention policy.
Avoid promising a fixed percentage saving from routing, caching, quantization, or prompt compression. The realized saving depends on traffic and quality.
Self-Hosting Economics
There is no universal query-count or GPU-utilization threshold at which self-hosting becomes cheaper. Compare equal-quality systems under the same latency, availability, safety, and governance requirements.
Capacity model
A first approximation is:
Here the first rate is measured during active serving and the duty fraction is the portion of wall-clock time serving this workload. If your rate is already averaged over the full interval, do not multiply by duty fraction again. GPU utilization is not the same as useful token throughput.
Benchmark the exact model, quantization, prompt-length distribution, output length, batch size, tensor parallelism, and serving stack. Prefill and decode have different bottlenecks; one average tokens-per-second number can hide tail latency and concurrency failures.
Reserved, serverless and interruptible capacity
| Capacity contract | Economic benefit | Required check |
|---|---|---|
| Reserved / committed | Predictable capacity or discounted commitment | Idle capacity, commitment length, failover headroom and cancellation terms |
| Serverless / autoscaled | Can reduce paid idle time | Minimums, cold starts, quotas, scaling speed and how active time is billed |
| Spot / interruptible | Potentially lower compute rate | Interruption probability, checkpoint/restart cost and deadline impact |
| Multiple providers / regions | More options for eligible workloads | Data placement, egress, model distribution, network latency and operational overhead |
No option has infinite burst capacity, and no universal 40% utilization crossover exists. Moving work to a cheaper region may violate residency requirements or cost more after egress and model loading. Use interruptible capacity for restartable work only when the recovery plan meets the deadline.
Memory floor: 70 billion parameters at two bytes each require about 140 GB (130.4 GiB) for weights alone. A nominal four-bit representation has a 35 GB (32.6 GiB) weight floor before scales, metadata and unquantized layers. Add KV cache, activations, runtime buffers and fragmentation. In a mixture-of-experts model, fewer active parameters per token do not mean only those weights must be stored. Benchmark the complete deployment; see quantization.
Full self-hosted TCO
Include:
- accelerator and CPU hours, including idle, failover, and capacity headroom;
- storage, networking, load balancing, orchestration, and observability;
- inference-server and model licenses;
- quantization or fine-tuning work and quality regression testing;
- security patching, abuse controls, incident response, and on-call labor;
- deployment, autoscaling, upgrades, and model migrations;
- downtime, capacity shortages, and disaster recovery.
A defensible break-even condition is:
Use current quotes for the intended hardware, provider, region, commitment, and network path. A public on-demand GPU price from another region is not a reliable production estimate.
Self-hosting can still be preferable when weight access, customization, data placement, or operational control matters more than the raw cost comparison.
Total Cost of Ownership and FinOps
FinOps is an operating practice for making technology spending accountable to business value through collaboration between engineering, finance and business teams. It includes allocation, forecasting and optimization; reducing spend alone is not the goal. FinOps Foundation framework.
Track the economic unit the product actually sells or operates:
Useful dimensions include tenant, feature, environment, model and version, prompt version, service tier, region, cache status, tool, success result, and retry/fallback reason.
Operational controls should include:
- budgets and alerts by tenant, feature, model, and environment;
- per-request usage and estimated-cost logs reconciled with invoices;
- anomaly detection for token growth, cache misses, loops, and retry storms;
- quotas and graceful degradation instead of one global hard stop;
- canary and shadow-traffic budgets;
- a dated rate-card registry with an owner and review cadence;
- invoice reconciliation for rounding, free allowances, taxes, and negotiated terms.
Forecast from the joint workload distribution, preserving relationships between long inputs, output length, tool use and retries. Report P50/P95 for planning, but do not multiply independent P95 values and call the result the P95 bill. Segment totals and replayed scenarios give a more defensible forecast. Run sensitivity analysis for price changes, growth, cache-hit rate, model migration, and quality regression.
Manager interview practice
Mental arithmetic: 2,000 input tokens at an illustrative $2/MTok plus 500 billed output tokens at $10/MTok costs $0.004 + $0.005 = $0.009 per call. Five such calls cost $0.045 before tools, retries, and review. These are invented rates to practice units, not a provider quote.
Recall checks: Are cache writes being counted twice? Does output already include reasoning? Which context/service tier applies? Are failed attempts in the numerator? What happens to unit economics if review doubles?
Define the successful-outcome denominator with product and the accounting boundary with finance. Maintain a rate version and forecast range rather than a permanent hard-coded model price.
Hypothetical API versus self-hosted TCO worksheet
Assume one million completed tasks/month of equal accepted quality. For teaching arithmetic only, suppose the API costs $0.008/task for model calls and the self-hosted serving fleet can support this workload at its measured latency target. These are invented rates, not current vendor prices.
| Monthly cost | API option | Self-hosted option | Assumption |
|---|---|---|---|
| Model calls or active GPU fleet | $8,000 | $6,000 | API scales with tasks; reserved GPU capacity is largely fixed |
| Standby/redundancy | Included in API rate here | $3,000 | Spare capacity for a replica failure |
| CPU, storage, networking | $1,000 | $2,000 | Application and serving overhead |
| Engineering/on-call allocation | $3,000 | $8,000 | Loaded labor allocated to this service |
| Evaluation and review | $2,000 | $2,000 | Same assumed accepted-quality target |
| Total | $14,000 | $21,000 | Migration cost excluded and reported separately |
| Cost/completed task | $0.014 | $0.021 | One million completed tasks, including unsuccessful-attempt cost in totals |
At 500,000 tasks, assume API model spend halves while its other lines stay fixed: $10,000 total, or $0.020/task. If the self-hosted fleet cannot shrink, $21,000 becomes $0.042/task. At two million tasks, the API becomes $22,000. Self-hosting is $0.0105/task only if the existing fleet can actually handle that demand and failure headroom; otherwise add the required capacity. The apparent break-even under fixed non-model assumptions is (21,000 − 6,000) / 0.008 = 1.875 million tasks/month, conditional on that capacity.
Idle and reserved capacity are already paid for in the GPU lines; do not add “idle cost” a second time. Measure utilization to explain why cost per task rises at low volume. Add one-time migration, compliance, tooling, and retraining costs over an explicit amortization period when relevant. For unequal quality, compare total cost per verified outcome rather than equal raw request counts.
Forecast the workload and test the boundary cases
A rate change can affect the whole request. With no cache or other fees, GPT-6 Sol at 272,000 input tokens and 1,000 billed output tokens costs 272000 × 2/1M + 1000 × 10/1M = USD 0.554. At 272,001 input tokens, the long-context rates apply: 272001 × 4/1M + 1000 × 15/1M = USD 1.103004. One added token crosses a billing boundary; do not estimate this as one extra token at the old rate.
For caching, use the strict inequality in the earlier section. At illustrative rates B=2, W=2.50, R=0.20 per MTok, a reusable one-MTok prefix costs:
| Total uses within the eligible lifetime | Without caching | One write, remaining uses hit | Saving |
|---|---|---|---|
| 1 | USD 2.00 | USD 2.50 | −USD 0.50 |
| 2 | USD 4.00 | USD 2.70 | USD 1.30 |
| 3 | USD 6.00 | USD 2.90 | USD 3.10 |
| 5 | USD 10.00 | USD 3.30 | USD 6.70 |
These are invented rates for isolating cache arithmetic, not a one-MTok short-context provider quote. Add storage and rewrite costs when they apply. If B ≤ R and W ≥ B, caching cannot reduce this token bill; the break-even expression requiring B > R is not applicable.
Traffic mix example: suppose 80% of tasks cost USD 0.005 and 20% cost USD 0.050, including their own calls and tools. Mean variable cost is 0.8 × .005 + 0.2 × .050 = USD 0.014, so 100,000 tasks cost USD 1,400. If the expensive slice rises to 40%, cost becomes USD 2,300, an increase of about 64.3%, with no rate or volume change. Preserve this correlation in the forecast.
Interview: design a cost-control service
Scope: build metering, budgets and allocation for an application using several model APIs. This is an internal cost ledger, not a payment processor. The provider’s invoice remains the reconciliation authority for its charges.
Functional requirements
- Attribute each model/tool execution to an authenticated tenant, product feature, environment and task.
- Select a versioned rate contract by provider, model, service, region and effective time.
- Reserve budget before admitting work and settle against observed usage.
- Include retries, failed tasks, storage, evaluations and human-review allocations in reports.
- Reconcile delayed usage and invoices without silently rewriting history.
Non-functional requirements
- Concurrent workers cannot each spend the same unreserved balance.
- A duplicated usage event cannot charge the internal ledger twice; distinct provider attempts remain distinct.
- Missing usage stays unknown and visible; it does not become zero after a deadline.
- Metering records exclude prompt bodies and credentials; access is tenant-scoped.
- Target p95 admission overhead below 20 ms and 99.9% admission-service availability at the illustrative peak below; validate both under hot-tenant load. Define bounded queues and behavior during ledger outages.
Initial design: each worker reads a tenant balance from a cache, calls the model, then subtracts an estimate. This is easy to prototype. Two workers can both read the same balance, a crash can lose usage, and a stream can end before final counters arrive. A nightly spreadsheet discovers overruns only after spending occurs.
Detailed design
Read diagram source
flowchart TB
U["Authenticated task and bounded execution plan"] --> A["Admission service: tenant budget + policy"]
R[("Versioned rates and billing contracts")] --> A
A --> B[("Atomic reservation and task ledger")]
B -->|Reservation committed| E["Execute bounded model / tool attempt"]
E --> P["Provider / external tool"]
P --> M["Usage adapter: disjoint units + actual tier"]
M --> Q["Durable usage event stream"]
Q --> S["Idempotent settlement / adjustment"]
S --> B
E -->|Outcome or usage unknown| H["Hold reservation; reconcile with provider"]
H --> S
I["Invoice / usage export"] --> C["Reconciliation and discrepancy review"]
C --> S
B --> D["Tenant allocation, forecasts and alerts"]
D --> O["Reviewed policy / routing change"]
O --> A
Data and API contract
| Record / operation | Required identity and behavior |
|---|---|
POST /reservations |
Tenant from authentication; unique operation ID, planned maximum exposure, currency, budget period and rate revision |
| Execution attempt | Task ID plus unique attempt ID; model revision, requested/actual tier, provider request ID and outcome |
| Usage event | Unique event identity, unit/category, quantity, observation time and whether cumulative or incremental |
| Settlement | Consume the matching reservation once; charge recorded actual usage and release only the known excess |
| Adjustment | Append a reasoned correction linked to the earlier entry and provider evidence |
| Cost report | Include settled amounts, outstanding reservations, unknown usage and allocation coverage separately |
Concurrency example: the limit is USD 10, settled spend USD 9.50 and existing reservations USD 0.10. Available balance is USD 0.40. Two workers each ask for USD 0.30. An atomic conditional update admits one; the other must be rejected or delayed. A read-then-write cache cannot enforce this invariant. Use a transaction or equivalent atomic conditional operation for the authoritative ledger. In a multi-region design, route a tenant budget to its authority or allocate bounded regional budgets; asynchronous replication alone can overspend a shared limit.
Settlement example: reserve USD 0.30, then observe USD 0.18 actual usage. Settle USD 0.18 and release USD 0.12. If the request outcome is unknown, retain the reservation until evidence resolves it. A lease expiry establishes worker ownership, not that the provider did no work. Stop further expensive actions when evidence is missing or the next action cannot be funded. Refunds or billing corrections are new adjustment entries, not deletion of an inconvenient charge.
An estimate is a strict spend cap only when every possible charge is bounded by the execution contract. Enforce maximum input/output, loops, tools, duration and approved services, and reserve conservatively for unknown cache hits. External pricing, delayed records or unbounded tools can invalidate that guarantee. Model cancellation can reduce further work; it does not undo completed billing. Keep a separate documented response for an actual charge exceeding the reservation.
Capacity: assume 100 admitted tasks/second at peak and three model attempts/task. This creates 300 attempts/second, before retries or shadow work. At least one reservation and one settlement mutation per attempt gives roughly 600 ledger mutations/second, plus events and reconciliation. Size partitions and indexes from the busiest tenant, not just total traffic. At nine million monthly attempts and an illustrative one-kilobyte usage record, raw records are roughly nine GB/month before replicas, indexes and event envelopes. Keep prompt traces under a separate privacy/retention policy.
| Flaw or change | Benefit of the repair | Cost / remaining tradeoff |
|---|---|---|
| Independent workers overspend | Atomic reservation protects a shared budget | Added latency and contention for a hot tenant |
| Every token stream update is summed | Normalize cumulative counters; settle final or corrected usage once | Provider-specific adapters and retained revisions |
| Duplicate event vs duplicate execution confused | Deduplicate event identity, preserve real attempts | More records; cannot erase external charges |
| Old requests use today’s price | Pin effective contract and append invoice adjustments | Rate registry and audit work |
| Cache estimate assumes every call hits | Reserve worst permitted uncached exposure, settle actual categories | Temporarily lower usable budget |
| Budget ledger is unavailable | Pause new expensive work or use preallocated bounded allowances | Availability tradeoff; no unlimited bypass |
| Global cutoff affects all customers | Tenant budgets and prioritized graceful degradation | More policy and operational complexity |
| Spending drops because quality drops | Join cost with verified outcomes and reviewer effort | Delayed labels and sampling cost |
Full cost and benefit calculation
For a smaller product example, assume 100,000 tasks/month. These rates and workloads are invented:
| Monthly item | Baseline | Proposed routing and cost controls |
|---|---|---|
| Model calls | 3/task × USD .009 = USD 2,700 | 2/task × USD .006 = USD 1,200 |
| Tools | USD .005/task = USD 500 | USD .004/task = USD 400 |
| Human review, 3 minutes at USD 60/hour | 5% of tasks = USD 15,000 | 6% of tasks = USD 18,000 |
| Common application operations | USD 2,000 | USD 2,000 |
| Incremental metering operations | — | USD 500 |
| Setup amortized over six months | — | USD 3,000 / 6 = USD 500 |
| Total | USD 20,200 | USD 22,600 |
Read diagram source
xychart-beta
title "Illustrative monthly total cost"
x-axis [Baseline, Proposal]
y-axis "USD per month" 0 --> 25000
bar [20200, 22600]
The model/tool bill falls 50%, yet TCO rises USD 2,400/month because reviewer work and new operating costs outweigh the saving. If both workflows produce 95,000 verified accepted outcomes after review, cost per 1,000 accepted outcomes rises from about USD 212.63 to USD 237.89. Equal accepted quality is an assumption to verify, not a consequence of human review.
Holding other lines fixed, the proposal breaks even at 5,200 reviewed tasks, or a 5.2% review rate. At 5%, it costs USD 19,600 and saves USD 600/month. That narrow margin needs production validation. The meter may still be justified for attribution and budget control even when a particular model-routing change fails its economic test; evaluate those decisions separately.
Closing remarks: “I would launch the meter with accurate attempt identity, atomic reservations and explicit unknown usage, then reconcile it with invoices. I would approve routing changes only after matched-quality tests show a reduction in full cost per outcome. Our example shows why token savings alone are insufficient. I would watch reviewer workload, hot-tenant admission latency and unresolved charges, with an owner for each discrepancy and a reversible routing policy.”
Interview questions with developed answers
Q1: How would you optimize a high-volume RAG application's cost?
Show answer and follow-up
Sample answer: I first attribute spending to ingestion, retrieval, generation, tools, retries, and review. I inspect calls per task, input and output lengths, cache categories, model mix, and successful outcomes. Then I test the largest plausible lever: removing redundant calls, improving evidence selection, using a cheaper acceptable configuration, or batching delay-tolerant work. I validate correctness, freshness, and permissions after each change. Caching is valuable only when its reuse and contract fit the workload. I would report net savings per successful task, not just a lower advertised token rate.
Follow-up: Why not shorten every answer? Missing explanations can increase user confusion and human rework.
Q2: When would you recommend self-hosting instead of APIs?
Show answer and follow-up
Sample answer: I compare equivalent quality, latency, availability, security, and data requirements. Self-hosting cost includes utilized and idle capacity, redundancy, serving software, staff, upgrades, and incident response. APIs include usage, tools, service modifiers, and application operating costs. Stable demand and control requirements may favor self-hosting, but there is no universal request-count crossover. I would benchmark a realistic traffic mix and run sensitivity analysis for utilization and future prices before recommending the investment.
Follow-up: What if the open model needs more retries? Include those attempts and their quality impact in the comparison.
Q3: How do you avoid double-counting cache and reasoning tokens?
Show answer and follow-up
Sample answer: I read the provider's usage and billing definitions, then partition tokens into non-overlapping billed categories. If total input includes cached input, I subtract the cached portion before applying the uncached rate. If reasoning is included in billed output, I do not add it a second time. I keep separate charges such as cache storage or tools explicit and reconcile estimates with actual invoices. Different APIs expose different fields, so a generic formula needs a provider-specific mapping.
Follow-up: Can the same visible answer cost different amounts? Yes, because input, reasoning, retries, service tier, and caching can differ.
Q4: Why can a cheaper per-token model be more expensive overall?
Show answer and follow-up
Sample answer: It may require more context, more calls, longer output, more retries, or more human escalation to achieve the same acceptable outcome. It may also miss difficult cases confidently, adding downstream correction costs. I compare complete workflows on matched tasks and calculate total cost divided by verified successful outcomes. I inspect severe failures separately rather than treating harm as a price tradeoff. A lower unit rate is one factor in that calculation, not the result.
Follow-up: What if review cost is paid by another team? It still belongs in an agreed whole-workflow comparison.
Q5: How do you budget when provider prices change?
Show answer and follow-up
Sample answer: I version the rate card and separate current rates, announced changes, and uncertain assumptions. I forecast baseline and adverse scenarios for prices, traffic mix, cache reuse, and escalation. I keep enough portability to evaluate alternatives without assuming they are behaviorally identical. Commitments or self-hosting decisions include the risk that demand changes or API prices fall. For a decision today, I verify the official terms and actual contract rather than rely on a memorized price from an interview guide.
Follow-up: What do you do with an unverifiable current rate? Label it as unverified and exclude it from a claimed current comparison.
Q6: The API reports 10,000 input tokens including 8,000 cached tokens. What is billed?
Show answer and follow-up
Apply the uncached rate to 2,000 and the cache-read rate to 8,000, under that stated usage contract. At USD 2/MTok and USD .20/MTok, input costs USD .0056. Add 1,000 output tokens at USD 10/MTok and the total is USD .0156. Cache writes/storage, if separately charged, remain explicit.
Follow-up: What if another provider’s input field excludes cache creation? Use that provider’s documented partition; do not copy this subtraction blindly.
Q7: One extra prompt token nearly doubles an example bill. Is that a calculation bug?
Show answer and follow-up
It can be a whole-request pricing threshold. The GPT-6 Sol example goes from USD .554 at 272,000 input tokens to USD 1.103004 at 272,001 with the same 1,000 billed output tokens. Apply the correct tier to the full request, then any eligible service/region modifiers.
Follow-up: Should you always trim below the threshold? Only if the removed context is unnecessary or an alternative architecture preserves required quality.
Q8: A prefix is used once before expiry. Does a 90% read discount save money?
Show answer and follow-up
There may be no cache read at all. At a 1.25× write rate, the one-use bill is higher than ordinary input. Model actual reuse, prefix churn, TTL and storage. The discount applies to eligible hit tokens, not all input or all task cost.
Follow-up: Can a semantic answer cache replace this calculation? It is a different mechanism with its own correctness, authorization and invalidation requirements.
Q9: Why can two small requests exceed a tenant’s remaining budget?
Show answer and follow-up
If workers independently read the same balance and only subtract after completion, both can spend it. Reserve exposure atomically before dispatch, then settle actual usage and release the known excess. Two USD .30 requests cannot both fit in a shared USD .40 balance.
Follow-up: Can asynchronous cross-region counters enforce one strict global cap? Not without coordination or bounded budget allocations that account for all regional spending.
Q10: A stream disconnects without final usage. What cost should the report show?
Show answer and follow-up
Show the known usage and the unresolved portion; retain a conservative reservation and reconcile using provider records. The provider may have generated output or charged tools after the client disconnected. Releasing the reservation on a timeout can let repeated uncertain attempts exhaust the real account.
Follow-up: Is this a free retry? No. A new execution can incur another charge even if our result storage deduplicates it.
Q11: A usage stream reports cumulative counts 100, 250 and 400. Is the total 750?
Show answer and follow-up
No, if these are snapshots of the same category and attempt: final cumulative usage is 400. Incremental deltas would be a different contract. Preserve event identity and revision, select or derive counters according to the provider schema, and avoid counting final totals on top of already settled increments.
Follow-up: What if final usage corrects an earlier amount? Append a linked adjustment with evidence rather than rewriting historical records silently.
Q12: Model and tool spending falls by half. Why might finance reject the optimization?
Show answer and follow-up
The worked comparison reduces those lines from USD 3,200 to USD 1,600, but review rises from USD 15,000 to USD 18,000 and new operations/setup add USD 1,000 monthly. Full cost rises from USD 20,200 to USD 22,600. Compare matched accepted outcomes and the entire agreed cost boundary.
Follow-up: At what review rate does this proposal break even? 5.2%, holding the other example assumptions fixed. That is a scenario result, not a general threshold.
Q13: When do nine million usage records require more than nine GB?
Show answer and follow-up
Nine million one-kilobyte records are about nine decimal GB of raw payload. Replication, indexes, envelopes, metadata and retained revisions add storage. Prompt bodies and traces may be far larger and should have their own access and retention rules. Estimate ingestion and queries as well as bytes.
Follow-up: Can storage be multiplied by every model call? Count the actual storage-time quantity once; do not repeatedly add the same monthly allocation.
Q14: A serverless GPU is cheaper per hour. Why can a reserved fleet still win?
Show answer and follow-up
The contracts differ. Compare billable active time, minimums, cold starts, required availability, scaling quotas and sustained demand. A reserved fleet can be efficient at steady load but costly when idle; serverless can still need warm capacity. Include retries and restart costs for interruptible instances.
Follow-up: Does a low-parameter MoE guarantee a small memory footprint? No. Active compute per token and resident weight memory are different quantities.
Q15: How would you forecast a bill after the traffic mix changes?
Show answer and follow-up
Replay or segment the joint workload distribution with the appropriate dated rates. In the example, increasing the USD .050 task slice from 20% to 40% raises 100,000-task cost from USD 1,400 to USD 2,300 with unchanged prices. Inspect correlated input/output lengths, tools and retries. Separately model announced price changes and uncertain assumptions.
Follow-up: Is multiplying every P95 input a P95 total forecast? No. Marginal percentiles do not determine a joint percentile or the monthly mean.
Final recall table and notes
| Decision | Evidence to bring |
|---|---|
| Estimate a task | All calls and tools, disjoint billed units, actual tier and dated rates |
| Approve caching | Reuse distribution, lifetime, writes/storage, permission and freshness rules |
| Enforce a budget | Atomic reservations, bounded exposure, reconciliation and unknown outcomes |
| Switch a model | Matched quality, total cost per accepted outcome, rollback criteria |
| Self-host | Measured capacity at the SLO, full memory, redundancy and staffed operations |
| Sign a commitment | Demand scenarios, contract terms, migration cost and price-change risk |
Interview tip: state units aloud, show one calculation and then challenge its assumptions. Keep quoted prices separate from workload assumptions. A rate table answers “what is charged”; a forecast answers “what will this workload cost”; a TCO comparison answers “which acceptable system is worth operating.”
60-second interview answer
I price the whole task, including every model turn, cache category, tool call, retry, and human review. I use the rate card for the actual model, context tier, service tier, region, and contract, then reconcile estimates with billed usage. I compare cost per successful outcome at the required quality and latency. The cheapest token price can lose if it needs more calls or creates more rework. I forecast a range and name the assumptions that would change the decision.
Remember: Rate × measured usage + tools + rework + operations.
Official Sources
Next checks
| Date / trigger | Recheck |
|---|---|
| Nov. 21, 2026 | GPT-5.6 Sol promotional guarantee; no automatic increase is announced |
| Jan. 1, 2027 | Announced Gemini Flash price and cache-storage increases |
| Every release or contract change | Model IDs, availability, rates, regions, tiers, free allowances and examples |
| Each invoice | Reconcile provider-metered usage and contract adjustments with internal estimates |
Text and specialist entries above were checked September 24, 2026:
- OpenAI API pricing
- Anthropic Claude API pricing
- Google Gemini Developer API pricing
- xAI API pricing
- Mistral API pricing
- DeepSeek models and pricing
Next: Model Selection Guide