Learnastra AI SYSTEM DESIGNAnup Rai

Complete design interview

Design an Adaptive AI Learning Tutor

By Anup Rai8 min readReviewed September 2026

Interview problem: build a tutor that explains approved course material, gives hints, checks answers and recommends the next exercise. Optimize learning and independent problem solving rather than conversation length.

This is a hypothetical Learnastra-style product exercise, not a description of Learnastra's deployed architecture. Volumes, targets and costs are assumptions. Start with adult learners and one bounded subject; educational use involving children adds separate consent, privacy and safeguarding requirements.

1. Clarify the learning contract

Ask what the learner should be able to do without assistance, which curriculum is authoritative, what counts as acceptable feedback, and whether assessments are practice or consequential grading. Assume formative interview practice; the system does not make hiring or certification decisions.

Functional requirements

  1. Explain concepts from a versioned, instructor-reviewed curriculum with source links.
  2. Offer graduated hints before a worked solution when the learner requests practice.
  3. Check free-text and structured answers against an explicit rubric; allow correction and disagreement.
  4. Track attempted skills and schedule review exercises with user control.
  5. Recommend exercises by prerequisites, observed gaps and the learner's stated goals.
  6. Let learners inspect, edit or delete retained profile information and export progress.

Non-functional requirements

  1. Target p95 first useful feedback below two seconds and completed ordinary feedback below eight seconds.
  2. Target 99.9% monthly practice-session availability, with static lessons and answer keys available during model outages.
  3. Measure delayed independent performance, feedback correctness and harmful/unsupported instruction separately.
  4. Isolate learner records and make deletion propagate to derived memory and analytics according to the retention contract.
  5. Bound calls, output lengths and weekly learner spend; avoid unlimited conversational loops.

2. Estimate the workload

Assume 20,000 daily learners, ten tutor turns each: 200,000 turns/day, about 2.31/s averaged over a day. A 20× peak is about 46.3/s. At a six-second mean, approximately 278 requests are in flight. If each turn uses 1,500 input and 250 output tokens, daily totals are 300M input and 50M output tokens.

A hypothetical $1/M input and $5/M output price gives $300 + $250 = $550/day for one model call per turn. Separate grading, hint generation and retries can multiply that. Compare complete practice-session cost and independently solved exercises, not tokens alone.

3. Build a small baseline

Architecture / visual model
flowchart LR L[Learner] --> API[Authenticated tutoring API] API --> C[(Reviewed lessons and exercises)] API --> M[Model generates a bounded hint] M --> CHECK[Rubric and source checks] CHECK --> L API --> P[(Exercise attempts)]
Read diagram source
flowchart LR
 L[Learner] --> API[Authenticated tutoring API]
 API --> C[(Reviewed lessons and exercises)]
 API --> M[Model generates a bounded hint]
 M --> CHECK[Rubric and source checks]
 CHECK --> L
 API --> P[(Exercise attempts)]

Start with authored exercises and answer rubrics. A deterministic checker can grade a numeric answer or validated code test; a model may provide a supplementary explanation. Preserve uncertainty for open-ended answers.

4. Find flaws and evolve deliberately

Flaw Repair Benefit Cost or limit
Tutor reveals answers immediately Explicit hint stages and learner-controlled reveal Supports effort before disclosure Some learners want faster answers
Fluent feedback marks a correct alternative wrong Rubric examples, alternative-answer tests and appeal Fairer, more reliable feedback Author/reviewer effort
Recommendation rewards easy completion Delayed unaided assessment and difficulty-aware reporting Better evidence of learning Longer experiments
Conversation summary invents a learning preference Provenance, confidence and editable memory Prevents stale inferred facts driving the path More state lifecycle work
A learner pastes instructions into an answer Treat answer as untrusted data, enforce tools in code Reduces authority confusion Does not guarantee perfect model behavior

5. Detailed architecture

Architecture / visual model
flowchart TD L[Learner web or mobile app] --> A[Identity and session API] A --> O[Bounded lesson orchestrator] O --> CUR[(Versioned curriculum and prerequisite graph)] O --> RET[Authorized lesson retrieval] RET --> CUR O --> H[Hint or explanation generator] O --> G[Deterministic or rubric-based grader] H --> V[Grounding and response checks] G --> V V --> L A --> E[(Append-only attempt events)] E --> S[Skill evidence projection] S --> REC[Prerequisite-aware recommendation] REC --> O A --> MEM[(Editable learner goals and preferences)] MEM --> O E --> EV[Offline evaluation and delayed learning experiments] EV --> RELEASE[Versioned release gate] RELEASE --> O
Read diagram source
flowchart TD
 L[Learner web or mobile app] --> A[Identity and session API]
 A --> O[Bounded lesson orchestrator]
 O --> CUR[(Versioned curriculum and prerequisite graph)]
 O --> RET[Authorized lesson retrieval]
 RET --> CUR
 O --> H[Hint or explanation generator]
 O --> G[Deterministic or rubric-based grader]
 H --> V[Grounding and response checks]
 G --> V
 V --> L
 A --> E[(Append-only attempt events)]
 E --> S[Skill evidence projection]
 S --> REC[Prerequisite-aware recommendation]
 REC --> O
 A --> MEM[(Editable learner goals and preferences)]
 MEM --> O
 E --> EV[Offline evaluation and delayed learning experiments]
 EV --> RELEASE[Versioned release gate]
 RELEASE --> O

Keep observed evidence separate from inferred skill estimates. An incorrect answer can indicate a misconception, a typo, an ambiguous question or bad grading; it is not automatically proof of low ability.

6. API and data design

POST /sessions chooses subject and goals. POST /sessions/{id}/turns accepts an idempotency key, exercise version and learner answer. POST /attempts/{id}/appeals records disputed feedback. DELETE /learners/me/memory/{id} revokes a retained preference.

Record Fields Rule
Exercise ID, version, skill IDs, prerequisites, prompt, rubric, author approval Grade against the attempted version
Attempt learner ID, exercise version, answer reference, hint level, grade version Do not erase the distinction between assisted and unassisted work
Skill estimate learner, skill, supporting attempt IDs, confidence, updated time Derived estimate, never an immutable fact
Memory source, user-confirmed flag, expiry, permission scope Retrieved preferences are data, not executable instructions

Store sensitive free text separately from minimal attempt metadata, with appropriate access and retention. User deletion creates a tombstone used by retrieval and projection workers so delayed events do not recreate removed memory.

7. Trace a practice session

  1. Authenticate the learner and load their explicit goal and current exercise state.
  2. Pick an exercise from prerequisite-eligible material; explain why it was suggested.
  3. Record the answer and assistance level before grading.
  4. Use a deterministic checker where valid; otherwise apply the pinned rubric with calibrated uncertainty.
  5. Generate bounded feedback from the rubric and authorized lesson evidence.
  6. Validate references and return feedback, a next hint or a clear inability to judge.
  7. Append the event, update skill evidence and schedule a future independent check.

Two concurrent devices may submit different answers. Use attempt IDs and optimistic version checks; do not let the latest arrival silently replace the first attempt. A model timeout does not erase the submitted answer.

8. Failure tests and evaluation

Test Expected behavior
Model outage Preserve answer; show authored hint/key or retry state
Grader disagreement Mark uncertain; retain both evidence and rubric; enable review
Curriculum update mid-session Finish against pinned version or explicitly restart
Deleted preference arrives from a delayed queue Tombstone/version check rejects resurrection
Learner repeatedly asks for final answers Respect the selected study mode; report assisted performance honestly
Prompt injection in an uploaded exercise No new tool authority; source treated as untrusted input

Evaluate feedback correctness on reviewed answer variants, groundedness, age/subject suitability for the supported audience, and accessibility. For learning outcomes, compare matched or randomized groups using delayed unaided tasks; account for baseline ability and attrition. Clicks and time in chat are engagement, not proof of learning.

9. Decisions and cost-benefit

Choice Benefit Cost or risk
Authored question bank first Known quality and reviewable rubrics Limited breadth and authoring cost
Generated variations More practice diversity Must validate correctness and difficulty
Simple skill evidence counters Explainable and easy to correct Coarse personalization
Probabilistic knowledge tracing Can model uncertainty and forgetting Data, calibration and explanation burden
Human feedback review Catches ambiguous grading Staffing delay and cost

Launch one subject with a reviewed set, compare to static worked examples, then expand. At $550/day model cost and 20,000 daily learners, the model line is $0.0275/learner-day before other calls, storage and staff. A cheaper model that causes more incorrect feedback is not a saving under the quality requirement.

10. Interview questions

Q1: How would you measure whether personalization works?

Sample answer: Use delayed independent exercises aligned to the intended skills, not the same examples seen during tutoring. Compare outcomes under a controlled study, record hint usage, account for prior ability and attrition, and review incorrect feedback separately from recommendation quality.

Q2: Should every wrong answer reduce a mastery score?

Sample answer: No. Preserve the observation and its context first: exercise difficulty, assistance, grader uncertainty and whether the prompt was ambiguous. Update an estimate with those limitations and let a learner dispute the result.

Q3: Why not use the whole chat history forever?

Sample answer: It adds cost, irrelevant or stale context and privacy exposure. Keep recent task state plus selected, attributable preferences under retention and deletion controls. A compact summary remains fallible and should not become authoritative.

Closing remarks and recall notes

I would start with reviewed exercises, explicit hint stages, auditable feedback and learner-controlled state. Personalization should improve demonstrated understanding. The main tradeoff is adaptation breadth versus feedback reliability and the cost of maintaining high-quality curriculum.

Remember Evidence
Teach a defined skill Curriculum and rubric
Separate help from mastery Assistance recorded on attempts
Personalize with uncertainty Editable preferences and attributable estimates
Measure later Independent delayed performance

Tip: A tutor is a learning system. Show where the learner's progress is measured independently of the language model's enthusiasm.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Research Reading for AI System Design
NEXT LESSONDesign an AI Gateway and Model-Routing Service →

Explore the diagram