Learn the concept, build a defensible design, and explain the decision clearly. This guide combines technical lessons with interview questions, diagrams, quantitative examples and complete system-design walkthroughs.
It is part of Learnastra, led by Anup Rai. Anup's background spans more than two decades of engineering and platform leadership, including Goldman Sachs and Consumer Reports. The teaching emphasis is practical: understand the mechanism, test the failure, and connect engineering choices to the user outcome. Employer names describe his background and do not imply endorsement.
Choose your starting point
| Your immediate need | Start here | Produce before moving on |
|---|---|---|
| Prepare for an interview | Practice hub → question bank | A spoken answer and an honest gap list |
| Understand how models work | LLM internals → attention | A small worked calculation or runnable example |
| Design retrieval | RAG fundamentals → production RAG | Separate ingestion and authorized answering paths |
| Build an agent | Agent fundamentals → tools and MCP → durable execution | One verified action and its timeout recovery |
| Choose a model or runtime | Model selection → serving | A workload-specific comparison, not a universal ranking |
| Evaluate a product | Evaluation foundations → release-gated evaluation | A versioned rubric, test set and release decision |
| Review access and risk | Access control → governance | A concrete allowed/denied action matrix |
| Move into an AI role | Role transitions → learning resources | A project that demonstrates a missing capability |
| Look up a term or pattern | Glossary → pattern reference | A definition, example and limitation |
Practice a complete interview
- Clarify the outcome. Identify users, essential behavior and exclusions.
- Write requirements. Number functional requirements and measurable nonfunctional targets separately.
- Draw a working baseline. Show preparation, live requests, stored state and external effects.
- Find specific failures. Trace a stale record, overloaded queue, missing permission or uncertain write.
- Justify repairs. Compare quality, latency, capacity, complexity and full operating cost.
- Close with a decision. State the compromise, release evidence and what would change the design.
Read diagram source
flowchart LR
C[Learn a concept] --> E[Work an example]
E --> Q[Answer without notes]
Q --> D[Design under constraints]
D --> F[Change a requirement or inject a failure]
F --> R[Review evidence and gaps]
R --> C
The question bank contains 40 quick checks, 128 developed answers, five complete design scenarios and ten leadership prompts. The whiteboard chapter adds nine worked exercises. These are authored practice materials, not a claim that employers use an identical question list or scoring rubric.
Explore the technical curriculum
| Area | What to learn | Entry lesson |
|---|---|---|
| Foundations | Tokens, embeddings, Transformer computation and inference | Tokenization |
| Model landscape | Capabilities, deployment eligibility and dated costs | Taxonomy |
| Training and adaptation | Fine-tuning, LoRA, preference learning, distillation and verification rewards | Adaptation |
| Inference | KV state, batching, precision, serving and edge deployment | Inference fundamentals |
| Prompting and context | Instructions, evidence selection and output contracts | Context engineering |
| Retrieval and data | Chunking, hybrid/graph/late-interaction retrieval, evaluation and source changes | Retrieval fundamentals |
| Agents | Planning, orchestration, tools, approvals, recovery and bounded loops | Agent fundamentals |
| Memory and state | Working context, durable facts, corrections, deletion and caches | Memory architectures |
| Frameworks | Choose abstractions and maintain compatible versions | Framework selection |
| Documents | Parsing, visual evidence, extraction and review | Document intelligence |
| Infrastructure and operations | Gateways, deployment, usage accounting and budgets | AI gateways |
| Security | Trusted identity, data boundaries and permitted effects | LLM application security |
| Reliability and governance | Failure policies, human oversight and applicable obligations | Reliability patterns |
| Evaluation and observability | Outcomes, traces, datasets, judges and uncertainty | Evaluation foundations |
| Design patterns | Recurring mechanisms and when they fail | Pattern reference |
| Tool and computer agents | Action interfaces, GUI state and execution boundaries | Tool-use landscape |
| Voice and audio | Turn-taking, streaming, interruption and confirmed actions | Voice agents |
| Multimodal generation | Media jobs, model compatibility, provenance and review | Multimodal generation |
For deeper evaluation practice, use the Phoenix and Langfuse guide and LangWatch and Langfuse guide. The research reference connects selected research questions to experiments and implementation decisions.
Choose a design to rehearse
| Design family | Worked interviews |
|---|---|
| Knowledge and search | Enterprise RAG, real-time search, knowledge management, MCP knowledge agent |
| Conversation and support | Conversational agent, support automation, voice healthcare |
| Code and computer work | Code assistant, autonomous coding, computer-use production |
| Analysis and decisions | Financial analysis, moderation, recommendations, compliance support, fraud |
| Platforms and pipelines | Multi-tenant SaaS, document intelligence, tenant fine-tuning, evaluation CI/CD, distillation |
The numbers in an interview scenario are assumptions to reason with unless explicitly identified as sourced measurements. A target such as 99.9% availability, 70% automation or a particular cost reduction is not a reported result merely because it appears in a worksheet.
Keep the basic definitions straight
| Term | Standard meaning and boundary |
|---|---|
| AI system design | Designing the complete application around an AI capability: behavior, data, models, interfaces, constraints, operations and evaluation |
| RAG | Supplying retrieved external information to a generative model at inference time; retrieval can be application-controlled |
| Agent | A system that chooses some next actions from observations while pursuing a goal; autonomy still has an enforced scope |
| Workflow | Prescribed orchestration that may include branches, models, tools and human steps |
| Chatbot | A conversational interface; it may use a fixed workflow, an agent or neither |
| MCP | A protocol connecting hosts through clients to tool/resource/prompt servers; it does not replace business authorization |
| A2A | A protocol for task communication between independently operated agentic applications |
| Evaluation | Measuring defined behavior against evidence and acceptance criteria, with failures and uncertainty accounted for |
See the FAQ for fuller explanations. A conversation is not automatically single-turn, and adding a model call does not automatically make a workflow an agent.
Reading, updates and access
Use the reader's search and chapter outline to find a mechanism. Follow linked concepts when a prerequisite is unfamiliar, then return to the interview question. Keep your notes focused on the definition, the failure you missed and the decision you would change.
Model names, provider prices, SDKs and regulations are time-sensitive. Relevant chapters include review dates and primary references. Updates are reviewed before publication; an external announcement does not automatically rewrite a lesson. For a deployment or purchase decision, verify the exact current provider contract.
The system-design module covers the broader distributed-systems interview curriculum. Preview the learning experience and see module and bundle plans. Checkout and tutoring booking are presented according to their actual availability; a draft price is not an active purchase or booked session.
For corrections, use the editorial and feedback guide. Applicable third-party notices are retained in the notices file; access to the hosted learning service and rights in individual materials are separate questions.
Final summary and notes
| Remember | Demonstrate it |
|---|---|
| Read to understand | Define the term in ordinary language |
| Recall to learn | Answer before opening the explanation |
| Design to reason | Trace one request and one failure |
| Measure to decide | Compare outcomes and complete costs |
| Review to improve | Revisit the specific gap after a delay |
Start with one concept and one related question. Finish by changing a constraint and explaining why the design changes.