This is a hypothetical interview scenario. Traffic, thresholds, latency and costs are planning assumptions, not measured results. Provider-specific behavior has source links; the proposed service has its own explicit decision contract.
Interview focus: make a bounded risk decision using fresh evidence, distinguish fraud policy from payment authorization, and learn from delayed, incomplete outcomes.
60-second interview answer
I would build a synchronous risk service that returns an allow, reject or additional-verification decision within 100 milliseconds. It would combine current payment data, fresh activity counters, historical features, a compact model and explicit policy rules. Each decision records the evidence and versions that actually produced it. Issuer authorization and any customer challenge happen through the payment workflow afterward. Explanations and investigations run asynchronously. I would evaluate fraud loss, legitimate declines, challenge completion and latency together, using mature outcome cohorts. Missing features or a model timeout would invoke an approved degraded policy rather than an invented low-risk score.
Remember: Identify attempt → Read fresh signals → Apply policy → Record decision → Learn from outcomes.
Interview problem and decision boundary
A payment processor handles 10 million transactions daily. Design a fraud service with a 100 ms decision target and a target false-positive rate below 0.1% of legitimate transactions.
Clarify the processor's role and the payment lifecycle. An allow decision means this fraud service permits the payment workflow to continue. It is not authorization by the issuer, settlement, proof of a legitimate cardholder or a guarantee against a later dispute. Extra authentication and manual investigation can take much longer than 100 ms.
A supported step-up may invoke 3D Secure through the payment platform. Whether a challenge occurs and what follows depend on issuer and integration behavior. Do not invent a universal “hold the card authorization while an analyst reads it” operation. A merchant's authorized-but-not-captured order review is a different workflow. Stripe risk outcomes, 3D Secure.
Functional requirements
- Evaluate a uniquely identified payment attempt using trusted merchant/account context and transaction attributes.
- Apply mandatory risk policies to every attempt, then score eligible attempts using an evaluated model and features.
- Return allow, reject or a supported additional-verification/review instruction before the caller's deadline.
- Record the actual decision basis, policy/model versions, feature freshness and degraded-mode status.
- Provide appropriately scoped customer explanations and richer analyst evidence without exposing evasion details.
- Join later outcomes, evaluate candidate policies/models and release changes with rollback.
Nonfunctional requirements
- Latency: target p99 under 100 ms for the risk service, measured from accepted request to response at the specified region boundary.
- Availability: target 99.99% timely policy responses, reporting degraded responses separately from normal decisions.
- Quality: track fraud loss and legitimate rejection rate, plus challenge and customer impact; 0.1% is a target to validate, not a universal threshold.
- Freshness: attach event time and materialization time to features; recent-activity counters have stricter limits than historical profiles.
- Replay safety: a retry of the same attempt returns the recorded decision and does not increment attempt counters again.
- Security and audit: restrict payment data and explanations, use tokenized identifiers and retain each record category under its approved schedule.
Do not state that every fraud record must be retained for seven years. Confirm jurisdiction, contracts, record type and deletion obligations. If the interviewer chooses seven years, use it as a sizing assumption, not a blanket legal claim.
Capacity and latency budget
| Quantity | Calculation | Implication |
|---|---|---|
| Average load | 10M / 86,400 ≈ 116 decisions/s | Average traffic is modest relative to the daily headline |
| Tenfold peak | About 1,160 decisions/s | Size each failover destination for its assigned recovery load |
| Concurrent requests | 1,160/s × 50 ms mean = 58 | This is Little's-law mean concurrency, not a p99 capacity guarantee |
| Decision records | 10M × 2 KB = 20 GB/day | About 600 GB per 30-day month before indexes and replicas |
| Seven-year archive assumption | 20 GB × 365 × 7 ≈ 51.1 TB | Excludes leap days, raw features, logs and redundant storage |
A compact model is often a useful baseline: for example, a locally served gradient-boosted tree model with a small feature vector. An ensemble adds calls, memory and operational risk; use it only when measured benefit warrants the extra work.
| Stage | Planning allowance |
|---|---|
| Request validation and routing | 10 ms |
| Parallel feature reads and freshness checks | 25 ms |
| Rules and compact model inference | 15 ms |
| Decision persistence | 15 ms |
| Response transport | 10 ms |
| Contingency | 25 ms |
| Total budget | 100 ms |
These are budgets, not independent percentile measurements. Measure the complete path under peak load, failover and dependency timeouts. Set dependency deadlines early enough to run the fallback and persist its result before the overall deadline. Do not place a remote LLM call on this path.
Start with a baseline that can explain itself
Begin with a versioned policy engine, a few reliable features, a compact trained model and durable decision records. Use reviewed reason templates for customer-facing messages. Keep the risk service stateless except for its dependencies, with local model artifacts and warmed connections.
Rules express explicit policy; models estimate patterns from data. A rule does not automatically make a model decision explainable. Record which mechanism actually determined the outcome. Run mandatory rules before optional model thresholds so a low score cannot bypass a blocked merchant or another mandatory control.
Detailed design
Read diagram source
flowchart TB
CALL[Payment orchestrator] --> API[Authenticate validate and deduplicate attempt]
API --> FEAT[Bounded parallel feature reads]
FEAT --> RULE[Versioned mandatory policy]
RULE --> GATE{Mandatory action?}
GATE -->|Yes| DECIDE[Resolve supported action and reason codes]
GATE -->|No| READY{Required features available and fresh?}
READY -->|Yes| MODEL[Local compact fraud model]
MODEL --> DECIDE
READY -->|No| FALLBACK[Approved degraded policy]
MODEL -->|Deadline or invalid score| FALLBACK
FALLBACK --> DECIDE
DECIDE --> STORE[(Decision record and outbox)]
STORE --> RESP[Return before deadline]
RESP --> CALL
STORE --> EVENT[Asynchronous decision events]
EVENT --> EXPLAIN[Templates or evidence-constrained explanation]
EVENT --> ANALYST[Scoped investigation queue]
OUTCOME[Issuer payment dispute and review outcomes] --> LABEL[(Outcome history)]
EVENT --> LABEL
Missing-feature status still passes through the mandatory policy stage. A missing input cannot silently disable a required control; its explicit unavailable-input policy applies. The fallback retains applicable mandatory restrictions.
The caller owns payment state. It maps the risk instruction to operations supported by the issuer/network and integration. If verification is unsupported or its deadline expires, use the explicitly approved alternate action. A response must never imply a challenge or hold happened when it only requested one.
APIs and records
| Operation | Contract |
|---|---|
POST /risk/decisions |
Attempt ID, immutable request hash and caller deadline; same attempt with changed fields is a conflict |
GET /risk/decisions/{id} |
Return the recorded result to an authorized payment orchestrator |
| Outcome event ingestion | Deduplicate source event IDs and append revised labels without overwriting history |
| Policy/model publication | Publish an immutable approved bundle with activation time and rollback target |
| Record | Fields and purpose |
|---|---|
| Payment attempt | Tokenized payment/account references, merchant, amount/currency, attempt ID, request hash and timestamp |
| Feature snapshot | Values, missingness, event cutoff, materialization timestamps and feature-definition version |
| Decision | Action, reason codes, model score/version, policy version, deadline result, degraded status and request hash |
| Outcome | Source event, payment status, fraud/legitimate/unknown label, observation time, effective time and adjudication state |
| Release bundle | Model artifact hash, feature schema, preprocessing, calibration, thresholds and signed-off evaluation |
Store sufficient evidence for the approved replay/audit purpose without indiscriminately logging full card data or raw identifiers. An immutable decision should identify the feature values actually read; replaying today's feature store does not reproduce yesterday's result.
Feature engineering and consistency
Precompute slow history, but update recent activity continuously. Useful signals include account tenure, historical spend distributions, recent attempts, known device associations and merchant-specific behavior. IP geography is approximate evidence, not proof of the cardholder's location.
Read diagram source
flowchart LR
HIST[Historical confirmed outcomes] --> BATCH[Versioned profile pipeline]
BATCH --> PROFILES[(Historical feature store)]
EVENTS[Deduplicated payment attempts] --> STREAM[Event-time and processing-time aggregates]
STREAM --> FAST[(Recent activity counters)]
CURRENT[Current request attributes] --> JOIN[Feature vector with timestamps and missingness]
PROFILES --> JOIN
FAST --> JOIN
JOIN --> CHECK[Freshness and schema checks]
CHECK --> SCORE[Model and policy]
Define a velocity feature precisely: “Distinct attempts for this account in the preceding 60 seconds, excluding the current attempt” differs from “all requests received in the current clock minute.” Count declined attempts when the policy is intended to detect card testing; approval-only counts can hide an attack.
- Assign the attempt ID before risk evaluation and deduplicate transport retries.
- Serialize or atomically update account-scoped recent activity so concurrent attempts cannot all see an empty history.
- Specify whether the returned feature includes the current attempt and use the same definition in training.
- Record event time, processing time and permitted lateness; late events update future evidence but do not rewrite past decisions.
- Treat account, device and merchant aggregates as separately versioned signals; cross-shard reads do not form one global atomic snapshot.
A Redis script can atomically deduplicate, update and read keys within the supported shard/key-slot boundary. It does not create global atomicity across unrelated shards or regions. Route an account to a home region or accept and quantify replication lag; add merchant/global attack signals with their own freshness limits. A hot merchant counter may require partitioned aggregation, trading immediacy for capacity.
An unusual spike across many cards can indicate coordinated fraud, but a flash sale can produce the same count. Combine merchant history, distinct-card/device patterns, failed verification and investigation evidence. “100 transactions in a minute” is not a defensible universal block rule.
Scores, thresholds and actions
A risk score is not necessarily a calibrated fraud probability. Calibration asks whether, among comparable cases assigned probability 0.1, roughly 10% are eventually confirmed fraudulent under the specified label definition. Validate this by cohort and monitor it over time. Do not compare numerical thresholds across unrelated providers or model versions.
The function below illustrates boundary behavior. Its 0.3/0.7 values are teaching cutoffs, not production recommendations or a provider's scale. ALLOW means only that this risk layer permits processing. Mandatory actions come from trusted policy code; the payment adapter verifies whether the requested action is supported.
import math
def decide(fraud_score, mandatory_action=None):
if mandatory_action is not None:
if mandatory_action not in {"REJECT", "HOLD", "STEP_UP"}:
raise ValueError("Invalid mandatory policy action")
return mandatory_action, "mandatory_policy"
if (isinstance(fraud_score, bool)
or not isinstance(fraud_score, (int, float))):
return "DEGRADED_POLICY", "model_unavailable"
if fraud_score < 0 or fraud_score > 1:
return "DEGRADED_POLICY", "invalid_score"
if not math.isfinite(fraud_score):
return "DEGRADED_POLICY", "model_unavailable"
if fraud_score < 0.3:
return "ALLOW", "low_model_risk"
if fraud_score > 0.7:
return "REJECT", "high_model_risk"
return "STEP_UP", "uncertain_model_risk"
In a deliberately simplified two-action example, suppose allowing fraud loses $60 and wrongly rejecting a legitimate payment costs $2. With a calibrated fraud probability p, allow has expected cost 60p and reject has expected cost 2(1−p). Reject becomes cheaper above 2 / 62 ≈ 3.23%. This calculation excludes authentication, recovery, margin, customer lifetime effects and regulatory constraints, but demonstrates why arbitrary 30%/70% cutoffs cannot substitute for policy economics.
Production selection compares the expected cost of all supported actions, subject to constraints. Challenge completion, authentication fees, remaining fraud after challenge and user abandonment all matter. Manual review has a finite queue and cannot be the destination for every uncertain score.
Explain the decision that actually happened
| Audience | Appropriate evidence | Avoid |
|---|---|---|
| Cardholder | Approved reason category and recovery action | Exact anti-fraud thresholds, secret signals or unsupported accusations |
| Investigator | Recorded features, fired rules, model contribution details and linked attempts | Treating feature attribution as proof of real-world causation |
| Auditor or internal reviewer | Versioned decision basis, policy approval and replay evidence | A fluent explanation generated from today's different policy |
Use templates first. If an LLM improves readability, give it only approved reason codes and permitted evidence after the decision. Validate its output and fall back to the template if it adds an unsupported cause. An issuer decline must not be explained as this model's fraud rejection.
SHAP can attribute a model prediction relative to its specified background and output representation. Contributions may be in raw score/log-odds units rather than probability points, depending on configuration. It does not establish causal fraud reasons, and a later language model cannot convert attribution into that proof. SHAP TreeExplainer.
Delayed labels and safe learning
Fraud may be discovered weeks after a transaction. Absence of a dispute today is not a confirmed legitimate label. Declines also create selective labels: a blocked transaction often has no observed outcome showing what would have happened if allowed. Training on the model's own reject decisions as “fraud” would reinforce its mistakes.
Read diagram source
flowchart TB
DEC[Versioned decisions and feature snapshots] --> JOIN[Join outcome history by attempt and payment]
LABEL[Disputes reviews and verified outcomes] --> JOIN
JOIN --> COHORT[Separate mature immature and unknown cohorts]
COHORT --> INVEST[Investigate drift and failure slices]
INVEST --> CAND[Candidate rule or trained model]
CAND --> OFF[Time-based offline evaluation and replay]
OFF --> SHADOW[Shadow score without changing decisions]
SHADOW --> APPROVAL[Risk owner approves bounded rollout]
APPROVAL --> CANARY[Canary with rollback and limits]
CANARY --> PROD[Versioned production bundle]
Build training features as they were available at decision time. An event timestamp alone may be insufficient if a fact arrived later: track availability/materialization time to avoid leaking future chargebacks into historical feature rows. Feature-store point-in-time joins are a useful mechanism, but still require correct data timestamps and label construction. Feast point-in-time joins.
Distinguish data drift (input distribution changes), concept drift (the relationship between signals and fraud changes) and pipeline faults (features suddenly become null). A distribution alert alone does not prove degraded fraud performance. Investigate timely proxy metrics while waiting for mature labels.
Emergency rules need an owner, bounded scope, expiry, replay tests, staged activation and rollback. Do not automatically loosen thresholds when legitimate rejection rises: an attack, label delay or feature outage may be responsible. Retrain when evidence supports it; a weekly schedule is an operational choice, not a guarantee of adaptation.
Evaluate the tradeoff with the right denominators
Assume a mature, adjudicated set of 10 million attempts has 0.2% fraud: 20,000 fraudulent and 9,980,000 legitimate. For a simplified binary reject/allow policy with 80% fraud recall and 0.1% false-positive rate:
| Outcome | Count | Meaning |
|---|---|---|
| True positives | 16,000 | Fraudulent attempts rejected |
| False negatives | 4,000 | Fraudulent attempts allowed |
| False positives | 9,980 | Legitimate attempts rejected |
| True negatives | 9,970,020 | Legitimate attempts allowed |
| Reject precision | 16,000 / 25,980 ≈ 61.6% | Nearly two in five rejections are legitimate despite the small false-positive rate |
This example assumes complete labels for illustration; production estimates must disclose selective labels and adjudication uncertainty. Report precision-recall behavior at the working threshold, dollar-weighted loss, approval rate and group/cohort breakdowns. ROC-AUC alone can hide poor performance at a very low false-positive operating point.
Measure customer-level impact separately: one affected customer can have many declined attempts. Also report the fraction of legitimate customers challenged, challenge completion, false-positive review burden and queue age. Step-up is neither automatically a false positive nor automatically a recovered sale; evaluate its actual outcome.
Outages, replay and operations
| Failure | Response | Tradeoff |
|---|---|---|
| Model misses its sub-deadline | Use approved fallback rules and explicit degraded action | Higher uncertainty and possibly more friction |
| Critical recent features are stale | Apply feature-specific fallback; do not fill missing risk with zero | Reduced acceptance or extra verification may be necessary |
| Decision persistence is unavailable | Use the approved audit/deadline failure policy; do not return an unrecorded normal result | Availability and audit guarantees must be chosen explicitly |
| Retry after response loss | Return the same attempt's durable decision | Request hash prevents reusing identity for changed payment data |
| Regional failover | Route to ready capacity and enforce replication-lag limits | Cross-region counters can lag and need conservative handling |
| Review queue exceeds capacity | Bound admissions and apply supported alternate policy | Humans cannot absorb unlimited uncertainty |
Reserve the caller-scoped attempt ID with a unique constraint and request hash. Concurrent retries join the same in-flight attempt; a fenced worker commits one winning decision. Counter updates also deduplicate the attempt ID so a crash between feature accounting and decision persistence cannot double-count the retry.
If persistence succeeds and the response is lost, the caller reconciles using the attempt ID. If the payment orchestrator creates a new attempt, that is a distinct event with a link to the previous attempt, not a retry that silently reuses its fraud decision. Repeated attempts may themselves be a risk signal.
Use authenticated release bundles for model and policy changes. Roll back the complete feature/model/calibration/threshold contract, not just a model filename. Monitor timely-response rate, normal versus degraded response mix, feature ages, duplicate attempts, rule firing, challenge outcomes and delayed fraud cohorts.
Operating costs and benefits
At a 1% explanation rate, 100,000 explanations daily using 500 input and 100 output tokens on GPT-6 Luna cost $10/day, or $300 per 30-day month at $0.10 input / $0.50 output per million tokens. Templates may make most of these calls unnecessary. OpenAI pricing.
| Item | Illustrative monthly amount |
|---|---|
| Feature stores, streaming and serving allowance | $8,000 |
| Archive, monitoring and operational allocation | $4,000 |
| Optional model explanations | $300 |
| 1,000 reviews/day × 8 minutes × $40/hour × 30 days | $160,000 |
| Partial operating cost | $172,300 |
The review queue in this example accepts only 0.01% of daily transactions. At 120 productive review hours per person per month it needs at least 34 reviewer equivalents; this is separate from engineering and fraud-loss cost.
Under the binary outcome example, 4,000 allowed fraud attempts at $60 average unrecovered loss imply $240,000/day. The 9,980 legitimate rejections at the simplified $2 impact imply another $19,960/day. These are modeled decision costs, not proven losses or interchangeable accounting categories. The dominant business tradeoff can dwarf the model-serving bill.
Interview follow-ups
1. Why not use an LLM for the risk decision?
The tight tail-latency and reliability budget favors a compact evaluated scorer and explicit policy. An LLM may help analysts or summarize approved evidence later, but must earn any new role through latency, accuracy and failure evaluation.
2. Can rules explain every machine-learning rejection?
No. A rule that did not fire is not the reason for the decision. Record the actual model and policy path, and use appropriate attribution with its limitations when the model drove the outcome.
3. What happens if the fraud service says allow but the issuer declines?
The payment remains declined. The risk service and issuer authorization are different decision-makers. Preserve both outcomes and their provenance so customer explanations and model labels remain accurate.
4. How do you stop concurrent card-testing attempts?
Deduplicate transport retries, atomically account for distinct attempts within the chosen key scope, and combine account/device/merchant signals. Explain the cross-shard and regional freshness limits; “a global counter” alone does not solve consistency or hot keys.
5. Why are recent chargeback rates misleading?
Recent cohorts may not have had time to produce disputes. Compare cohorts at similar maturity and keep unknown labels separate. Also account for blocked transactions whose counterfactual fraud outcome is unobserved.
6. Can we automatically approve cheap payments during an outage?
Only if an explicitly approved degraded policy permits a bounded class with monitored exposure. An unconditional low-value bypass can be exploited at scale. The interview answer should name the policy owner, limits and recovery behavior.
7. How do you select a production threshold?
Use representative mature cohorts, calibrated scores where probabilities are needed, action-specific costs and constraints. Evaluate legitimate rejection, challenge outcomes, fraud loss and subgroup impact; test the whole policy before a bounded rollout.
Closing notes
The design is a deadline-bound risk decision with traceable evidence, embedded in a longer payment lifecycle. Freshness, retries, supported actions and delayed labels matter as much as model accuracy. Close by stating the selected latency/availability tradeoff, the expected customer and loss impact, and the exact fallback and rollback paths.
Related: Evaluation frameworks, Reliability patterns.