Interview problem: automate routine e-commerce support across 12 languages, integrate with existing help-desk systems, and transfer sensitive or unresolved work to people without duplicate actions or false promises.
An answer is a message. A resolution is an outcome that addresses the customer's issue under a defined measurement rule. A shipped-order lookup, an issued refund and a handoff are different outcomes; sending a fluent reply does not establish any of them.
This is a hypothetical Learnastra interview scenario with 2M incoming tickets/month and a proposed 60% automation objective. Workload, costs and staffing are assumptions. The business outcome is correct resolution with acceptable customer effort, not maximizing automation at any cost.
1. Requirements and success criteria
Clarify which actions may be automatic, who approves refunds, how customer identity is established across channels, what a request for a human means operationally, and which system owns conversation state. Identify the order/payment system of record separately from Zendesk or Salesforce ticket status.
Functional requirements
- Ingest ticket/message events from supported channels and existing Zendesk/Salesforce workflows.
- Authenticate the customer before exposing account or order details.
- Route between automatic handling, assisted handling and specialist/human handling.
- Retrieve current order facts and applicable policy; execute only authorized actions.
- Support human takeover during a conversation and suppress obsolete bot work.
- Reply in supported languages with quality checks and a language-capable fallback.
- Track actions, evidence, recontacts, customer feedback and verified resolution status.
Nonfunctional requirements
- Operate continuously across the agreed 12 languages and channel mix.
- Propose p95 first useful response below three seconds for routine online messages; durable background actions have separate completion targets.
- Propose 99.9% monthly availability for message intake, with explicit degraded behavior when order/payment services fail.
- Deny cross-customer access and prevent unauthorized or duplicate financial actions.
- Preserve conversation/action state across duplicate webhooks, out-of-order updates and worker failures.
- Bound model, tool, queue and retry costs; define whether the five-cent target covers automation only or the whole support operation.
- Measure quality and queue delay by language, issue and route, alongside automation rate.
Resolution definition for this exercise: an issue has a confirmed appropriate outcome, no unresolved required action, and passes the chosen follow-up/recontact checks. A request for a human is honored according to the service policy; it is not a bot resolution merely because it was transferred.
2. Establish traffic and human capacity
With 2M tickets/month over 30 days, average intake is about 2M / 2,592,000 ≈ 0.772 tickets/second. At an assumed eight inbound/outbound processing events per ticket, average event load is about 6.17/s; a tenfold peak is about 61.7/s. Message count, tool calls and repeated classification determine inference load, not ticket count alone.
At 60% complete automation, 800,000 tickets/month still require people. If each needs ten minutes of active work and each staff member supplies 120 productive hours/month, capacity is 120 × 60 / 10 = 720 tickets/person/month. The workload requires about 1,112 productive staff equivalents, before extra coverage and specialized handling. A diagram that ends at “human queue” has not solved this capacity requirement.
3. Baseline: read-only order lookup with a handoff
Start with verified customer identity, an authorized order-status tool, a short evidence-grounded reply and a clear path to a person. Do not start by giving a general chat model permission to issue refunds.
| Baseline failure | Improvement | Benefit | Cost or limitation |
|---|---|---|---|
| Order ID belongs to someone else | Customer-scoped lookup | Prevents object-access failure | Authentication and ownership checks |
| Model treats estimate as guarantee | Typed tool result and claim validation | More accurate commitments | Source can still be delayed or wrong |
| Same webhook produces two replies | Durable deduplication and send intent | Repeat-safe processing | Provider-specific receipt reconciliation |
| Bot sends after takeover | Serialized ownership and dispatch reservations | Stops new obsolete actions | Already dispatched work still needs reconciliation |
| Refund times out after succeeding | Durable operation identity and unknown state | Prevents unsafe duplicate attempts | Reconciliation queue and staff |
| Fluent translation changes policy meaning | Per-language evaluation and specialist review | Better language coverage | Review and translation cost |
4. Detailed architecture and integration contracts
Read diagram source
flowchart TD
CH[Web, email and help-desk events] --> VERIFY[Verify source, deduplicate and persist intake]
VERIFY --> CONV[(Conversation, ownership and message log)]
CONV --> ROUTE[Policy routing and required evidence checks]
ROUTE -->|Eligible routine work| AUTO[Authorized read-only tools]
ROUTE -->|Draft or action needs review| ASSIST[Draft and exact proposal for reviewer]
ROUTE -->|Human requested or specialist need| HUMAN[Staffed human queue]
AUTO --> DRAFT[Grounded reply]
ASSIST --> HUMAN
HUMAN --> APPROVE[Approved reply or action with owner version]
APPROVE -->|Approved action| ACT[Policy executor and durable action intent]
APPROVE -->|Reply only| CHECK
ACT --> OMS[Order/payment system of record]
OMS --> RECEIPT[Confirmed or unknown outcome]
RECEIPT --> CHECK[Claim, privacy and language checks]
DRAFT --> CHECK
CHECK --> SEND[Ownership-aware dispatcher and send intent]
SEND --> CHOUT[Customer channel]
SEND --> LOG[(Delivery receipts and resolution evidence)]
RECEIPT -->|Uncertain| RECON[Reconciliation queue]
RECON --> HUMAN
Zendesk and Salesforce are integrations, not substitutes for internal durable state. Verify Zendesk webhook authenticity, persist accepted events before acknowledging them, and tolerate retries. Fetch the current ticket when an event is incomplete or stale. Avoid update loops by distinguishing the integration's own changes from genuinely new customer work.
Salesforce Pub/Sub event retention is 72 hours for platform/CDC events. Store replay checkpoints, but plan reconciliation after a longer gap. Replay IDs are opaque and not necessarily contiguous; do not increment them as ordinary sequence numbers or treat the stream as permanent history.
APIs and records
POST /conversations/{id}/messages
{client_message_id, text, language_hint?} → message_id
POST /conversations/{id}/takeover
{expected_owner_version} → new_owner_version
POST /actions/{id}/approve
{proposal_hash, owner_version} → approval_id
GET /actions/{id}
→ proposed|approved|dispatched|confirmed|rejected|unknown, receipt?
| Record | Essential fields |
|---|---|
| Conversation | Verified customer reference, channel mappings, owner type/ID, owner version and issue state |
| Message | Provider/channel ID, author, original text, language and delivery/processing state |
| Evidence | Order/policy version, retrieval time, permitted audience and typed facts |
| Action proposal | Stable operation ID, customer/order, action, amount/currency, policy version and proposal hash |
| Approval | Authorized actor, exact proposal, owner version and validity/revocation |
| Action attempt | Reserved dispatch, provider key, request digest, outcome and receipt/reconciliation state |
| Handoff | Reason, language/skills, queue deadline, summary, transcript and pending actions |
Channel identity is not automatically account identity. A sender's email address or a model-extracted customer ID cannot replace the required account-verification flow.
5. Route on policy first, then evaluated quality
| Route | Eligibility | Typical case | Success evidence |
|---|---|---|---|
| Automatic | Allowed task, verified identity, complete facts and passing checks | Current order status | Correct sourced reply and issue outcome |
| Assisted | Human judgment/approval needed, but useful evidence or draft available | Policy exception or refund proposal | Authorized review and confirmed action if any |
| Human/specialist | Explicit human request, restricted issue, serious risk or unresolved failure | Legal complaint or repeated unsuccessful attempts | Ownership transfer and staffed follow-through |
A refund may be automatic under an explicit narrow business policy in another product. The model's confidence cannot create that policy. VIP routing follows the contracted service tier, not a universal rule that all VIP questions are risky.
Read diagram source
flowchart LR
INPUT[Conversation and trusted state] --> HARD{Human required by request or policy?}
HARD -->|Yes| HUMAN[Human or specialist queue]
HARD -->|No| EVID{Identity, facts and required checks complete?}
EVID -->|No| CLARIFY[Clarify or assisted handling]
EVID -->|Yes| ELIG{Task eligible for automation?}
ELIG -->|No| ASSIST[Human-reviewed proposal]
ELIG -->|Yes| QUALITY{Evaluated quality gate passes?}
QUALITY -->|Yes| AUTO[Bounded automatic workflow]
QUALITY -->|No| ASSIST
Fit any quality threshold against labeled outcomes by language/task, with the cost of wrong automation and review load made explicit. Combining “angry,” “VIP” and a model's self-confidence with arbitrary weights does not yield a calibrated probability. Re-evaluate routing when new facts arrive; a routine order lookup can become a disputed-charge case.
6. Tools establish facts and constrain actions
This executable tool factory binds customer identity in application state. The model can request an order lookup but cannot supply a different customer ID:
def make_order_status_tool(authenticated_customer, oms_client):
def get_order_status(order_id: str) -> dict:
order = oms_client.get_order_for_customer(
order_id, authenticated_customer.id
)
return {
"status": order.status,
"shipped_date": order.shipped_at,
"estimated_delivery": order.eta,
"tracking_url": order.tracking_url,
}
return get_order_status
The data adapter must actually enforce ownership, handle not-found/forbidden without leaking another customer's existence, and return typed current values. Validate tracking-link destinations and distinguish stale or absent estimates. The model can still misread a correct result; tool use reduces unsupported answers but does not eliminate hallucination.
For a proposed refund, the executor validates customer/order ownership, remaining refundable amount, currency, policy, required approval and current action state. Use exact monetary units and trusted business rules. Approval for $40 does not authorize $50, another order or a different recipient.
Refund succeeds, but the response is lost
- Persist the exact action intent and stable business operation ID before dispatch.
- Atomically reserve a permitted dispatch under current conversation ownership and approval.
- Call the payment service with a stable idempotency key and immutable parameters.
- On confirmed success, store the provider receipt and update the support case.
- On timeout or ambiguous response, record unknown, block independent duplicate attempts and reconcile.
- Query the provider or retry the same operation only within its documented deduplication contract.
An idempotency key identifies repeat attempts at one operation; it does not make every failure retryable or reserve the identity forever. Stripe's idempotency contract, for example, describes key retention and cached responses. After a key can be forgotten, blindly retrying it may create a new action. Human staff must use the same action ledger instead of creating a second refund to “fix” the uncertain first one.
Tell the customer “the refund outcome is being confirmed” while it is unknown. Canceling the support job does not reverse a payment already issued. See durable execution.
7. Human takeover is a concurrency problem
An ownership label in the UI is insufficient. Increment a conversation ownership version atomically, and serialize that change with dispatch reservation for every bot reply/action. New attempts from the old version are rejected. A check followed by an unguarded send leaves a race.
Define the boundary clearly: work already reserved/dispatched before takeover may still complete. Surface it to the human and reconcile the outcome. If the product requires takeover confirmation only after all sends drain, make that waiting state explicit and bounded. Do not promise that an external message can always be recalled.
async def handoff_to_human(conversation_id, agent, ownership, summaries):
# Adapter verifies agent access and atomically transfers dispatch ownership.
conversation = await ownership.claim_for_human(conversation_id, agent)
try:
summary = await summaries.from_transcript(conversation.messages)
except Exception:
summary = "Summary unavailable; read the attached transcript."
return {
"owner_version": conversation.owner_version,
"summary": summary,
"full_transcript": conversation.messages,
"confirmed_actions": conversation.confirmed_actions,
"pending_or_unknown_actions": conversation.pending_actions,
}
Bound summary generation time so a stalled model cannot delay handoff. A generated summary is a convenience; the underlying transcript and receipts remain authoritative. Required handoff content includes the issue, attempted solutions, explicit user requests, unresolved commitments and pending/unknown actions.
Optional metadata—language, service tier, user-supplied accessibility needs and cautious sentiment labels—helps routing but does not grant authority. Prefer “asked twice for a person” over a confident psychological label. Preserve the origin/age of inferred fields and allow correction.
8. Multilingual response and pre-send checks
Read diagram source
flowchart LR
MSG[Original customer message] --> LANG[Language and script detection]
LANG --> NATIVE[Evaluated native-language processing]
LANG -->|Needed for a downstream step| TRAN[Controlled translation with original retained]
TRAN --> NATIVE
NATIVE --> TOOLS[Authorized facts and action outcomes]
TOOLS --> REPLY[Reply in the customer's supported language]
REPLY --> CHECK[Names, numbers, policy and meaning checks]
CHECK -->|Pass| SEND[Ownership-aware send]
CHECK -->|Uncertain or unsupported| REVIEW[Language-capable reviewer]
A single multilingual model can cover several languages, but coverage must be evaluated. Translation is optional, adds latency/cost and can alter a legal complaint, negation or currency. Preserve original text and test code-switching, transliteration, dates, policy terms and requests for a human. A detector's language guess is not the customer's identity or location.
Before sending:
- Verify factual claims and commitments against order/policy/action evidence.
- Keep estimated delivery distinct from a guarantee and a proposed refund distinct from a confirmed one.
- Enforce audience restrictions; internal notes and another customer's details must not appear.
- Check understandable language, appropriate tone and the company's actual communication policy.
- Revalidate current ownership and create a repeat-safe send intent.
A keyword filter cannot establish whether a promise is authorized. A competitor mention is not inherently unsafe. Retrieved policies and customer messages can contain prompt injection; trusted tools and application policy must enforce authority outside the prompt. See guardrail limits.
9. Economics: resolve the five-cent ambiguity
Assume these totals across all routine automation calls per incoming ticket, using GPT-6 Luna standard short-context uncached rates of $0.10 input/$0.50 output per million tokens. Billed output includes reasoning where applicable. API pricing.
| Component | Total input / output tokens | Cost per incoming ticket |
|---|---|---|
| Intent/routing classification | 1,000 / 100 | $0.00015 |
| Response generation | 4,000 / 600 | $0.00070 |
| Additional model checks | 1,000 / 100 | $0.00015 |
| Tools, retrieval and monitoring | Hypothetical allowance | $0.00300 |
| Automation subtotal | Before other exclusions | $0.00400 |
At 2M tickets this is $8,000/month. Allocating all of that spend to 1.2M automatic resolutions gives about $0.00667 per auto-resolution. That calculation excludes human labor, help-desk licenses, translation, deeper-model work and retries beyond the token assumptions.
At an assumed $5 human cost for each of 800,000 escalated tickets, add $4M/month. Even if all 2M tickets eventually resolve, the partial full-operation cost is (4M + 8,000) / 2M = $2.004 per resolution. Thus a five-cent target fits the illustrated automation component, but does not fit total support cost. Clarify the business requirement rather than silently changing its denominator.
With this simplified $8,000 automation budget, a $0.05 × 2M = $100,000 overall allowance leaves only $92,000 for human work, or 18,400 tickets at $5 each—less than 1% of arrivals, before omitted costs. This exposes why “60% automation” and “five cents full cost” are inconsistent under the given assumptions.
Model choice follows per-language resolution quality, latency and total cost. A small model can handle structured routine work; difficult drafts may justify a larger model under review. The cheapest token rate is not necessarily the cheapest correct resolution.
10. Measure outcomes and operate the service
Track verified task outcomes, unsupported commitments, unauthorized/duplicate actions, automation eligibility/coverage, handoff wait, unknown-action age, customer effort, satisfaction and recontact. Report denominators and issue mix by language and route.
A same-issue reply within 24 hours can signal failure, but may also be thanks, a clarification or a new event. Combine defined issue matching, customer confirmation, authoritative action status and sampled review. No reply can reflect abandonment rather than success. Detect repeated apologies/no progress and stop the loop with a useful handoff.
| Failure drill | Required behavior |
|---|---|
| Duplicate/reordered help-desk events | One durable message/action intent, current state preserved |
| Order service unavailable | State missing information; no fabricated ETA |
| Payment success followed by timeout | Unknown state and reconciliation, no independent duplicate |
| Human takeover races with bot reply | Enforced dispatch boundary and visible in-flight work |
| Approval revoked or amount changed | Reject stale/mismatched action |
| Customer asks for a human in another language | Correct route without an English-only keyword dependency |
| Human queue is full | Honest queue/contact expectations and incident ownership |
| Provider outage creates retry surge | Bounded retries, tenant/route budgets and protected capacity |
Pilot read-only order status, then reviewed drafts, then narrowly approved actions. Shadow and canary model/routing changes on held-out cases before broader release. Support operations owns staffing and handoff targets; payments operations owns unknown financial outcomes; platform engineering owns dispatch, deduplication and integration recovery.
Interview follow-ups
1. How do you route a confident legal complaint? Apply the specialist policy before the quality score. Confidence estimates answer reliability; it cannot override authority or the customer's request for a person.
2. Why is order ID alone insufficient for a lookup? It identifies an object, not the caller's permission. Bind the tool to verified customer identity and enforce ownership in the underlying data query.
3. What if the bot keeps apologizing? Detect missing progress and unresolved required actions, then route usefully. Politeness and response count are not evidence that the issue was solved.
4. What should happen after an ambiguous refund timeout? Keep the operation ID, report uncertainty and reconcile under the provider's contract. Both automation and staff use the same action ledger to avoid a second refund.
5. What exactly stops the bot when a person takes over? The ownership version is serialized with send/action dispatch reservation. Old work cannot reserve new effects; previously dispatched effects remain visible and require outcome handling.
6. Can one model serve all 12 languages? It can be a candidate, but test task and policy accuracy in each language. Use translation or trained reviewers where needed, preserving the source message and measuring the added errors and delay.
7. Is five cents per resolved ticket achievable here? The automation-only example fits it. With 40% human handling at $5 each, total cost exceeds $2 per resolution before additional overhead. The interviewer should see that arithmetic and the unresolved business tradeoff.
60-second interview answer
I would begin with verified read-only order support, then add actions only through a policy-controlled executor. Routing separates automatic, assisted and human handling, with explicit human requests and risk rules taking priority over confidence. Durable operation IDs and unknown states protect payment recovery, while serialized ownership prevents new bot actions after takeover. I would evaluate every supported language, staff the remaining queues, and report verified resolutions and full operating cost rather than counting sent replies as success.
Remember: Identify → Retrieve → Authorize → Confirm → Resolve or transfer.