FinOps is an operating discipline that connects technology spending to business value through shared engineering, product and finance accountability. For AI, the unit of work may include model calls, retrieval, tools, serving capacity and human review. A cheaper token is useful only if the complete workflow still produces an acceptable outcome.
The FinOps Foundation describes the broader discipline. This chapter applies it to an interview-learning service without assuming one universal AI margin, cache-hit rate or self-hosting break-even point.
Start with standard cost terms
| Term | Meaning | Example |
|---|---|---|
| Attribution | Connect usage to its source | A feedback request belongs to the coaching feature |
| Allocation | Assign shared costs using an agreed rule | Split shared serving cost by measured resource consumption |
| Showback | Report costs to accountable teams | Monthly feature-level cost dashboard |
| Chargeback | Apply allocated costs to internal budgets/accounts | Debit the coaching team's budget |
| Unit economics | Cost/value per defined unit | Cost per verified useful feedback session |
| Forecast | Estimate future usage/cost under assumptions | Baseline, growth and low-cache scenarios |
| Commitment | Spending/capacity obligation over a period | Reserved serving capacity paid even when idle |
Use allocation rules that teams can understand, and leave unallocated spending visible. Unit economics requires a clearly defined denominator. A conversation ending is not automatically a successful coaching outcome.
A complete model-call bill
Sum each non-overlapping billed quantity multiplied by its applicable rate. Providers differ in input, cache-read/write, output/reasoning, image/audio, tools, service tiers, context tiers and contractual terms.
Illustrative call, using invented rates:
| Quantity | Amount | Rate per million tokens | Charge |
|---|---|---|---|
| Uncached input | 1,000 | $2.00 | $0.0020 |
| Cached input | 3,000 | $0.20 | $0.0006 |
| Billed output, including any billed reasoning in this category | 1,000 | $10.00 | $0.0100 |
| Total model charge | $0.0126 |
If a provider's input_tokens includes cached input, charging all of it at the uncached rate and then adding cached input double-counts. Likewise, do not add reasoning tokens again when they are already included in billed output. Cache-write pricing may replace another input rate rather than being an additive fee. Current OpenAI caching documentation makes that distinction explicitly.
Keep raw provider usage, normalized billing categories, rate revision, currency and calculation version. Reconcile with invoices; delayed records, credits, discounts and rounding can make a live estimate differ from the final bill.
Move from calls to tasks
A user task can trigger several model calls, retries, tools, searches and subagents. Add all of them. Then include the non-model operating costs required by the metric you are reporting.
Suppose a monthly cohort contains 100,000 attempted coaching tasks:
model and tool charges, including failed attempts = $2,000
5,000 reviews × 2 minutes × $30/hour = $5,000
model + review subtotal = $7,000
verified successful tasks = 90,000
subtotal per success = $7,000 / 90,000 ≈ $0.0778
This deliberately excludes hosting, support and other costs, so label it a subtotal. Model cost per attempt is $0.02; it answers a different question. Show attempts, successful outcomes and cost together so excluding difficult tasks cannot make the system appear better.
If a cheaper model halves model/tool charges to $1,000 but doubles review demand to $10,000, the subtotal becomes $11,000. At the same 90,000 successes, cost per success rises to about $0.1222. The token bill improved; the workflow economics worsened.
For margin reporting, use the organization's agreed cost-of-revenue policy. Do not mix development investment and operating costs inconsistently, or assume all AI products share the same margin gap relative to other software.
Interview design: a cost and budget service
Functional requirements
- Attribute attempts to task, account, feature, model and release.
- Normalize usage and calculate versioned estimates.
- Reconcile usage with provider billing and handle corrections.
- Reserve spend before parallel work and enforce configured limits.
- Report actuals, forecast, unit costs and unexplained differences.
Non-functional requirements
- Avoid duplicate charges in accounting when usage events are retried.
- Prevent concurrent workers from each spending the same remaining budget.
- Keep private prompts out of routine cost records.
- Preserve an auditable history of rate changes and adjustments.
- Define behavior during delayed metering or an unavailable budget service.
Start with trustworthy usage records and showback. Add chargeback after identifiers and shared-cost rules are reliable; inaccurate incentives encourage teams to move cost into someone else's category.
Read diagram source
flowchart TD
A[Task with trusted account and feature] --> B[Atomic budget reservation]
B -->|Allowed| C[Bounded model and tool work]
B -->|Insufficient budget| D[Queue, reduce optional scope or stop]
C --> E[Attempt usage events]
E --> F[Deduplicate and normalize]
F --> G[Versioned cost ledger]
G --> H[Reconcile reservations and invoices]
H --> I[Actuals, forecast and unit economics]
I --> J[Product and engineering decisions]
J --> B
Budget controls that survive concurrency
Consider a task with $1.00 remaining. Two workers each want a $0.70 reservation. Independent reads of the balance let both proceed and commit $1.40. A shared atomic reservation must allow at most one under that budget.
- Assign a stable reservation/operation ID and applicable account/task limits.
- Atomically check and reserve a conservative amount before dispatch.
- Enforce output, tool, step, concurrency and time limits during execution.
- Reconcile actual usage and release only the unused reservation.
- Deduplicate repeated usage events and handle late corrections explicitly.
- Keep uncertain in-flight spending reserved until its outcome is resolved; do not free it merely because a worker lease expired.
Hard caps require enforceable upper bounds and complete metering. If a tool has an unknown charge or a provider reports usage late, call the control an estimate with a defined buffer—not an exact guarantee. On limit exhaustion, queue suitable work, omit optional work with a clear explanation, or return an explicit partial result. Do not silently skip correctness or access checks to save money.
Diagnose a cache regression with arithmetic
Assume 100,000 daily calls share an eligible 8,000-token prefix. There are 800 million eligible prefix tokens. At hypothetical rates of $0.20/million for hits and $2/million for misses, excluding output and cache-write/storage charges:
| Prefix hit rate | Hit tokens | Miss tokens | Daily prefix cost |
|---|---|---|---|
| 80% | 640 million | 160 million | $128 + $320 = $448 |
| 10% | 80 million | 720 million | $16 + $1,440 = $1,456 |
That is a $1,008/day increase with unchanged traffic. Investigate prompt/tool ordering, request-specific fields near the prefix start, model changes, expiry, eviction and the provider's matching rules. A percentage of requests with some hit is not the same as the fraction of input tokens reused.
Keep reusable content stable when semantics permit; place request-specific content afterward. Verify the resulting behavior and actual usage. Do not extend retention beyond policy just to improve a cost metric. Prefix caching reuses model processing; it still generates a new answer.
Response caching reuses a completed answer and needs access, freshness, policy and version checks. Exact query matching is not automatically safe if the user's identity or underlying data changed. Semantic caching adds false-match risk and must be evaluated for the task.
Choose a cost lever from the measured driver
| Driver | Candidate improvement | What to check before claiming a saving |
|---|---|---|
| Repeated stable context | Prefix caching | Eligible tokens, write/read/storage prices and actual reuse |
| Repeated valid answers | Response caching | Access, freshness and false-hit consequences |
| Overpowered model for a narrow task | Smaller model or routing | End-to-end quality, escalation, latency and review |
| Too many repair/retry calls | Fix error causes and bound retries | Completion rate and errors hidden by retrying |
| Unnecessary retrieved/history tokens | Better retrieval, context selection or compaction | Missing evidence and lost state |
| Verbose or unbounded output | Useful output contract and limits | Truncation, completeness and task quality |
| Repeated narrow high-volume task | Distillation or fine-tuning | Training/data/maintenance costs and held-out performance |
| Delay-tolerant bulk work | Batch processing | Model/feature support, deadline and partial-failure behavior |
| Steady serving demand | Capacity commitment or self-hosting | Utilization, redundant capacity, staffing and lock-in |
Do not assume retrieval is always cheaper than long context: repeated cached context and a small corpus can change the comparison. Do not assume distillation preserves every long-tail capability. Use the cost optimization playbook and distillation lesson for the technical tradeoffs.
Batch economics are useful only when the workflow can tolerate the actual completion contract and supported features. The OpenAI Batch API documents model support, a completion window and per-item results/errors. A batch can expire with unfinished work; reconcile completed items before retrying. “No person is watching” alone is not sufficient if another system needs the output sooner.
Forecast scenarios instead of one optimistic number
Model changes in these drivers separately:
- Active users and tasks per user.
- Task mix and calls per task, including retries/subagents.
- Input/output distributions and cache behavior.
- Model/provider/service-tier mix and rate changes.
- Review rate, time per review and peak staffing needs.
- Fixed commitments, spare capacity, storage and support.
- Evaluation jobs, backfills and onboarding spikes.
A reserved fleet costs money while idle. Pay-per-use can handle demand variability but has quotas and service limits. Forecast both low-utilization and growth cases before committing. Monthly reviewer hours do not prove that a team can meet a peak-hour review deadline.
For an optimization costing $6,000 to implement and saving an estimated net $1,000/month, simple payback is six months. Include ongoing maintenance and quality-related costs in “net.” Revisit after real results arrive; a forecast is not a booked saving.
Close the operating loop
The FinOps for AI overview connects cost visibility and controls to business outcomes. A practical monthly review should produce decisions:
| View | Decision it supports |
|---|---|
| Actual versus forecast | Revise assumptions or investigate anomalies |
| Cost per verified outcome plus quality | Change scope, architecture or model configuration |
| Feature/account allocation and gaps | Assign ownership and repair missing attribution |
| Review/rework and failed-attempt cost | Fix quality problems that token charts hide |
| Commitments and utilization | Adjust capacity or future procurement |
| Optimization proposals | Approve an owner, expected benefit and reassessment date |
A lower bill is not automatically success if fewer useful outcomes were delivered. Higher total spend can be healthy when valuable usage grows and the unit economics remain acceptable.
Interview questions and answer checks
- Traffic is flat but the bill doubled. Where do you start? Reconcile periods/rates, then compare task mix, calls, token categories, cache reuse, retries, tools and releases.
- Why is output text length a poor proxy for cost? Billed reasoning, tool calls, context and multiple attempts can dominate the visible answer.
- Can exact-match response caching return the wrong answer? Yes, when identity, permissions, source data or relevant configuration differ.
- Why reserve spend before parallel work? Independent balance checks can each authorize spending the same remaining amount.
- What if a worker crashes before reporting usage? Preserve the unresolved reservation and reconcile provider/operation records before releasing it.
- How do showback and chargeback differ? One reports allocated cost; the other applies it to internal budgets/accounts.
- Should every offline task use batch? Only when its deadline, feature support and recovery needs fit the batch contract.
- How do you choose an optimization? Use the measured cost driver, estimate net saving and implementation cost, and validate quality and cost per useful outcome afterward.
Final notes
Remember attribute → reconcile → forecast → control → improve value. Keep attempts, outcomes and costs visible together. The objective is a sustainable useful service, not the smallest token counter.