Start with the work you want to do, then connect your existing experience to evidence that you can do it. An “AI” title can describe product development, model research, infrastructure, evaluation, customer deployment or leadership. These roles share some concepts but require different depth.
This Learnastra guide provides eight practical preparation paths. They are study plans, not guaranteed hiring timelines or formal occupational categories. Reviewed September 24, 2026.
Understand the work before choosing the title
| Work family | Typical responsibility | Evidence worth preparing |
|---|---|---|
| Application engineering | Build useful features around models, data and business workflows. | A tested application with clear interfaces, failure handling and measured quality. |
| Model engineering and research | Train, adapt or investigate model behavior. | A reproducible experiment with a meaningful baseline and defensible conclusions. |
| Infrastructure and platform | Serve models, manage capacity and operate shared capabilities. | Load and cost analysis, rollout design and recovery from a failed dependency. |
| Evaluation and quality | Design datasets, graders and processes that measure behavior. | A rubric, error analysis, validated evaluator and uncertainty-aware comparison. |
| Security and reliability | Investigate threats, enforce boundaries and contain failures. | A threat model, tested control and incident or recovery walkthrough. |
| Product and program | Choose user outcomes, priorities, experiments and delivery constraints. | A decision brief connecting quality, adoption, cost and risk. |
| Engineering leadership | Develop teams, make architecture decisions and own delivery. | Truthful examples of technical judgment, delegation, coaching and accountability. |
| Customer deployment | Turn customer requirements into working integrations. | Discovery, implementation, rollout and a measurable customer outcome. |
Current employer examples illustrate the variation. OpenAI's Backend Software Engineer (Evals) combines backend systems with evaluation infrastructure. Its Frontier Evals & Environments research role emphasizes experiments and environments. Anthropic's Forward Deployed Engineer posting combines customer work and production implementation. These postings were checked on the review date; they are examples, not a market survey or a promise that the openings remain available.
Build a personal gap map
- Select two or three current roles whose responsibilities you actually want.
- Separate required skills from preferences and marketing language.
- Match each responsibility to a real project, experiment or demonstrated skill.
- Mark gaps as can explain, can implement with help, can implement independently, or have operated in practice.
- Choose the most consequential gap and one bounded project that can demonstrate progress.
- Confirm the interview format, expected implementation language and permitted tools.
Read diagram source
flowchart LR
J[Target responsibilities] --> E[Existing evidence]
E --> G[Prioritized gaps]
G --> P[Bounded project and study]
P --> V[Tests, measurements and review]
V --> I[Mock interview and explanation]
I --> R[Revise the gap map]
R --> G
Example: “I have built asynchronous APIs” is relevant to model integration. It does not yet show that you can handle a streamed partial answer, an ambiguous tool outcome or a model-quality regression. Prepare those specific additions rather than treating all previous experience as either irrelevant or sufficient.
1. Backend engineer to AI application engineering
Your API, database, concurrency and reliability experience transfers directly. Add the model-specific behavior and data boundaries that ordinary CRUD services may not expose.
| Existing skill | Addition to practice | Evidence |
|---|---|---|
| API contracts | Model output, tool-call and streaming contracts | Separate completed, partial, refused, failed and timed-out outcomes. |
| Databases and search | Retrieval, embeddings, indexing and source provenance | Trace a returned answer to authorized current evidence. |
| Authentication | Access across retrieval, memory, caches and tools | Demonstrate a cross-tenant request being rejected at the relevant boundary. |
| Async processing | Long-running jobs, cancellation and reconciliation | Recover an operation after the caller loses its response. |
| Caching | Prefix computation versus reusable application responses | Explain a safe key and an invalidation event. |
| Operations | Quality evaluation alongside latency and availability | Detect a wrong-answer regression while the service is still returning HTTP 200. |
Practice sequence
- Learn inference, structured generation and model selection. Build one bounded call with explicit error handling.
- Build a retrieval baseline over permitted documents. Include scope and provenance from the start. Compare lexical, vector and hybrid approaches only when the task justifies them.
- Add a tool workflow, durable operation identity and recovery. Run a small release comparison with quality, latency and complete costs.
Portfolio deliverable: a document-support application that handles missing evidence, changed permissions, deletion and a failed dependency. Include a requirements list, baseline diagram, detailed diagram, measured flaws and a justified repair. Use the enterprise RAG interview as a preparation reference.
Interview questions
- Why did you add vector retrieval, and what evidence would justify removing it?
- The external action succeeded but your API timed out. How do you prevent a duplicate effect?
- How can an answer be wrong when retrieval latency and availability are healthy?
2. Frontend engineer to AI product engineering
Streaming interfaces, accessibility, state management and user feedback are valuable foundations. Extend them to uncertain and partial model outcomes, with clear server-side authority.
| Existing skill | Addition to practice | Evidence |
|---|---|---|
| Rendering and state | Partial output, tool progress and reconnect behavior | The interface distinguishes a draft from a confirmed result. |
| Forms and validation | Structured proposals and user confirmation | The user approves the exact action, not a vague earlier intent. |
| Async error handling | Cancellation, late responses and retries | A cancelled view does not display a late result as current. |
| Accessibility | Keyboard navigation, reading order and restrained live announcements | Streaming does not repeatedly interrupt assistive technology. |
| Content display | Safe Markdown/HTML handling and source links | Untrusted generated content cannot execute code or impersonate trusted UI. |
| Analytics | Feedback tied to model, prompt and task versions | Ratings can be investigated without pretending they are unbiased ground truth. |
Practice sequence
- Build a streamed response interface with loading, partial, completed, cancelled and error states. Use a server boundary for provider credentials and authorization.
- Add cited suggestions that a user can inspect and accept or reject. Keep source navigation, keyboard focus and layout stable while content arrives.
- Instrument useful observations, validate feedback collection and test reconnects, duplicate events and stale responses. Compare UI behavior under slow and failed requests.
Portfolio deliverable: a document editor with reviewable AI suggestions, clear citations and recoverable state. Demonstrate one accessibility check and one malicious-output case as well as the happy path. Follow the code-assistant design for proposals, review and confirmed changes.
Interview questions
- Does stopping text generation also cancel an already-issued tool action?
- How do you stop old streamed events from overwriting a newer request?
- Why is thumbs-up rate insufficient to establish factual correctness?
3. QA engineer to evaluation and AI quality
Test design, exploratory investigation and regression prevention transfer well. AI evaluation adds sampled behavior, semantic judgments, annotation uncertainty and statistical interpretation. Some evaluation roles also require substantial backend or research skills.
| Existing skill | Addition to practice | Evidence |
|---|---|---|
| Test cases | Representative data, edge cases and distribution slices | Explain where cases came from and what population they represent. |
| Bug investigation | Trace-based error analysis and failure categorization | Distinguish retrieval, reasoning, permission and execution failures. |
| Test automation | Deterministic checks and semantic evaluators | Use a parser for structure and a validated rubric for meaning. |
| Acceptance testing | Expert labeling, disagreement and adjudication | Preserve ambiguous cases rather than silently forcing consensus. |
| Regression suites | Paired comparisons and repeated trials | Report changed outcomes, uncertainty and important regressions. |
| Release process | Quality gates, unknown outcomes and monitoring | Failed evaluator calls cannot disappear from the denominator. |
Practice sequence
- Read evaluation fundamentals. Inspect a small permission-safe set of traces and write specific observed failures before choosing metrics.
- Build a versioned dataset and rubric. Separate development from held-out validation. Compare a model judge with expert labels; measure false acceptance, false rejection and unjudged cases.
- Add a CI comparison and a reviewable release report. Set thresholds from the task's consequences and available evidence, rather than adopting an arbitrary “faithfulness above 0.85” rule.
Portfolio deliverable: an evaluation package that reproduces a baseline/candidate result, exposes regressions by category and includes an evaluator failure. Keep the environment, versions and permitted dataset reproducible. Use the evaluation study guide, implementation companion and release-gating interview.
Interview questions
- Your model judge agrees with reviewers 95% of the time. What does that fail to tell you?
- How would you evaluate a rare but severe failure without confusing an enriched test set with production prevalence?
- What should happen when the evaluator is unavailable during a release check?
For a security-focused role, add threat modeling, adversarial test design and runtime isolation. Quality testing alone is not evidence of red-team expertise.
4. Product manager to AI product management
User research, prioritization, experiments and clear requirements remain central. Add an explicit understanding of model uncertainty, evaluation costs and the difference between a convincing demonstration and dependable user value.
| Existing skill | Addition to practice | Evidence |
|---|---|---|
| Problem discovery | Decide whether AI improves the actual workflow | Compare with a non-AI or human-assisted baseline. |
| Success measures | Task quality, coverage, correction and escalation | Define success per eligible user task, not just generated responses. |
| Roadmaps | Data, evaluation and operating dependencies | Explain who owns labels, exceptions and source freshness. |
| Experiments | Model/version changes and noisy feedback | Separate causal evidence from a dashboard correlation. |
| Business cases | Cost per correct outcome and human capacity | Include review, rework, incidents and adoption assumptions. |
| Communication | Honest limits and safe user expectations | Define where the system should clarify, defer or decline. |
Practice sequence
- Learn the distinction between prompting, retrieval, adaptation and agents. Explain a relevant design without requiring a vendor name.
- Participate in expert review of outputs and traces you are authorized to see. Define the important failure categories, their consequences and the handling of uncertainty.
- Write a limited launch plan with measurable success, operating capacity, stop conditions and ownership. Evaluate adoption and task outcomes alongside model quality.
Portfolio deliverable: a product decision brief supported by a small prototype or evaluated sample. State functional requirements, quality constraints, excluded uses, evidence limitations and a complete business case. Use the support automation interview and full-cost pattern example.
Interview questions
- When is a higher abstention rate a product improvement, and when is it a failure?
- If an assistant reduces time per task but adoption is low, how do you revise the business case?
- Who decides whether a disputed answer is acceptable, and how is that judgment recorded?
5. Engineering manager to AI engineering leadership
Leadership experience transfers through hiring, coaching, prioritization and technical accountability. AI work adds particular model, data and evaluation dependencies; it does not replace ordinary engineering management or make delivery speed the only measure of success.
| Existing responsibility | Addition to practice | Evidence |
|---|---|---|
| Architecture decisions | Model/data/evaluation tradeoffs | A decision record that compares viable alternatives and states assumptions. |
| Definition of done | Behavioral quality and release evidence | The team knows which failures block release and who can accept residual risk. |
| Incident response | Quality, misuse and data incidents | A runbook can disable a feature or model route while preserving investigation evidence. |
| Staffing and coaching | Complementary engineering, domain and evaluation skills | Ownership is explicit rather than assigning all AI work to one specialist. |
| Delivery planning | Provider changes, data quality and review capacity | Dependencies have owners, fallback plans and realistic capacity estimates. |
| Team development | Honest evidence of learning and judgment | Coaching rewards finding important flaws, not only producing more generated code. |
Practice sequence
- Build enough depth to inspect a complete design and challenge its assumptions. Prioritize model economics, evaluation, security and the architecture relevant to your team.
- Define responsibilities for data, metrics, releases, incidents, customer communication and human escalation. Run a tabletop exercise for a quality regression with normal infrastructure health.
- Prepare examples of coaching, disagreement, prioritization and recovery from your actual career. Explain personal decisions and team contributions accurately, including what did not work.
Portfolio or interview evidence: a reviewed architecture, skills/ownership map, rollout plan and truthful leadership stories. A sanitized written case study or internal presentation can demonstrate this without exposing employer confidential information. Practice with behavioral preparation.
Interview questions
- The product team wants a release while the quality team identifies a severe regression. How do you structure the decision?
- What work should you delegate, and what accountability remains yours?
- How do you decide whether to buy a managed runtime or operate one internally?
A move to director or VP depends on organizational scope and demonstrated leadership. There is no reliable calendar that turns a study plan into that promotion.
6. Platform or DevOps engineer to AI infrastructure
Your experience with deployment, observability, capacity and failure recovery is directly relevant. Add model-serving resource behavior, quality-aware releases and the economics of hardware versus hosted services.
| Existing skill | Addition to practice | Evidence |
|---|---|---|
| Scheduling and capacity | GPU memory, batching, sequence state and token throughput | Size capacity from request lengths, concurrency and hardware measurements. |
| Autoscaling | Model loading, queue growth and warm capacity | Explain cold starts, drain behavior and admission limits. |
| Monitoring | First-token, inter-token and completion latency | Separate queueing and model execution from application overhead. |
| CI/CD | Versioned models, prompts, data and evaluators | Roll back a behavioral regression, not only a failed container. |
| Secret management | Tenant-scoped provider credentials and tool authority | Rotation and failover preserve access boundaries. |
| Cost management | Utilization, quotas, review and operations | Compare the full service at equivalent quality and availability. |
Practice sequence
- Study inference, KV state and batching. Select an actual supported model that fits the hardware and license; do not copy a nonexistent size/version from a tutorial.
- Exercise a small serving deployment, or begin with an API-backed queue if you lack suitable GPU access. Measure the actual route you run; do not present API measurements as self-hosted GPU results.
- Add admission limits, timeouts, failure injection and a quality-gated rollout. Compare hosted and self-operated costs at a defined workload and utilization level.
Portfolio deliverable: a serving/load report with request-length distribution, concurrency, latency percentiles, failures and full costs. Include a dependency outage and a rollback. Use AI infrastructure, CI/CD and FinOps.
Interview questions
- Why can larger batches increase throughput while making interactive latency worse?
- How does a tenfold increase in context length affect memory and capacity?
- What must remain compatible when requests fail over to another provider?
7. Data engineer to AI data and retrieval engineering
Data lineage, schemas, quality and incremental processing are strong foundations. AI adds document structure, training/evaluation boundaries and derived representations whose versions and permissions must remain aligned.
| Existing skill | Addition to practice | Evidence |
|---|---|---|
| ETL and parsing | Document structure, OCR and reading order | A table or heading remains interpretable after extraction. |
| Incremental pipelines | Chunking, embeddings and index generations | Reprocess changes without serving a mixture of incompatible versions. |
| Schema and lineage | Source IDs, offsets, provenance and permission scope | Trace every retrieved passage to an authorized source version. |
| Data quality | Semantic errors, duplicates and poor labels | Detect useful failure classes rather than relying on minimum chunk length alone. |
| Dataset management | Training, development and test separation | Prevent source or answer leakage across split boundaries. |
| Streaming updates | Corrections, deletion and revocation | Changes propagate through indexes, caches and retained summaries. |
Practice sequence
- Build ingestion over permitted PDF, HTML or document samples. Preserve source structure and versioned metadata. Read document processing and chunking.
- Build a representative retrieval evaluation set with explicit relevance judgments. Compare a small number of justified embedding candidates and an exact-match baseline. Investigate disagreement instead of automatically deleting low-agreement labels.
- Implement an index migration with rollout and rollback. Test updates, duplicate events, deletes, a changed access policy and a partial failure. Explain whether distribution drift changed task quality.
Portfolio deliverable: a versioned ingestion and retrieval pipeline with lineage, quality reports and a demonstrated deletion/migration procedure. If converting production traces into training examples, establish permission, redaction, retention and split rules first. Use AI data engineering, production RAG and the embedding migration interview.
Interview questions
- What happens when the embedding dimension or tokenizer changes?
- How do you delete one source from every derived representation that can expose it?
- Why can a shifted document distribution be harmless while unchanged aggregate metrics hide a serious regression?
8. Model engineering or research preparation
For roles that actually require training or research, application integration alone is insufficient. Conversely, a research role's requirements should not be imposed on every AI application job.
| Existing foundation | Addition to practice | Evidence |
|---|---|---|
| Programming and mathematics | Tensor shapes, gradients, objectives and numerical behavior | Explain and debug a small implementation. |
| ML experimentation | Baselines, controls, ablations and reproducibility | Separate the changed variable from confounding changes. |
| Model training | Data rights, split integrity, checkpoints and resource limits | Reproduce a run and identify its cost and failure conditions. |
| Statistical reasoning | Uncertainty, multiple comparisons and distribution shift | State what the result supports and what remains untested. |
| Reading papers | Assess definitions, methods and evidence | Explain a paper's mechanism without repeating its headline as a universal claim. |
Practice sequence
- Study the model internals and training basics needed for the target role.
- Reproduce a bounded result or adaptation experiment with an appropriate baseline. Use resources whose hardware and data requirements you can actually meet.
- Add an ablation or negative test and write a concise report with uncertainty, limitations and a defensible next experiment. Follow the research reading guide.
Portfolio deliverable: a reproducible experiment that demonstrates understanding. A result that disproves your hypothesis can be valuable if the method is sound. Confirm degree, publication and research-experience requirements in each actual posting; job titles do not settle them.
Interview questions
- How do you know the improvement came from your method rather than more compute or a changed dataset?
- What evidence would falsify your explanation of the result?
- Which reported metric is least representative of the intended deployment, and why?
A flexible twelve-week preparation structure
This is an optional planning example. Adjust it for your starting knowledge and available time; it does not predict when you will get hired. At six hours a week, twelve weeks provides 72 study hours, not twelve weeks of full-time engineering experience.
| Stage | Suggested weeks | Work | Exit criterion |
|---|---|---|---|
| Select and understand | 1–2 | Analyze target roles; review the essential concepts. | Explain the target work and identify the two most important gaps. |
| Build the baseline | 3–5 | Implement a bounded project and document its requirements. | Another person can run or inspect it and understand the request path. |
| Measure and repair | 6–8 | Investigate failures, compare one improvement and calculate costs. | The claimed improvement has reproducible evidence and known limitations. |
| Operate and explain | 9–10 | Test failure/recovery and prepare a design walkthrough. | Explain a rollback, access boundary and important operational tradeoff. |
| Rehearse and revise | 11–12 | Conduct mock interviews and close the highest-impact gaps. | Answer follow-ups without relying on memorized product names. |
Keep a weekly record with four fields: what I learned, what I built, what failed, what I will investigate next. If the project is not understood or reproducible, extend that stage instead of declaring it complete because the calendar ended.
Make a portfolio reviewable
- State the problem and scope. Include numbered functional requirements and non-functional constraints.
- Show the baseline. Explain why it was a reasonable starting point.
- Document the actual implementation. Identify model/service versions, data origin and important interfaces.
- Show failures. Include at least one meaningful edge case or recovery case.
- Explain the improvement. Present evidence, tradeoffs and any regression.
- Report complete economics. Separate measured spend, estimates, human capacity and hypothetical savings.
- Make it reproducible. Supply permitted sample data, setup steps and the checks that matter.
- Explain limits. Describe what you did not test and what production deployment would still require.
| Weak claim | Stronger, honest evidence |
|---|---|
| “I built a RAG app.” | “I compared lexical and hybrid retrieval on a held-out set and investigated the missed exact identifiers.” |
| “It is production-ready.” | “I tested these failure cases at this load; these dependencies and scale limits remain untested.” |
| “The judge is 95% accurate.” | “Here is the label distribution, confusion matrix, disagreement process and unjudged fraction.” |
| “AI saves 50%.” | “Under these workload and review assumptions, the full cost changes from this baseline to this candidate.” |
| “I led the whole project.” | “I owned these decisions; these colleagues owned the other components.” |
These are phrasing examples. Substitute your own work and results. Do not copy their implied experiences into a résumé.
Prepare for the actual interview
Use the question bank to find gaps and the whiteboard exercises for complete design practice. A useful rehearsal sequence is:
- Clarify the user's problem and excluded scope.
- State functional requirements, quality constraints and workload assumptions.
- Draw the basic request/data path.
- Identify the most consequential flaw under those assumptions.
- Improve the design, explain costs and test the change.
- Discuss state, permissions, failures, scaling and operations.
- Close with the decision, its largest uncertainty and the next validation step.
Match the employer's format: coding, architecture, research discussion, product case, customer exercise or management conversation. Use AI assistance only when permitted by that process. A take-home exercise and a live closed-tool interview can have different rules.
Tip: names are secondary to mechanisms. Being able to explain idempotency, retrieval recall, queueing, calibration or a safe approval flow is more useful than reciting a framework catalog without understanding its behavior.
Common preparation mistakes
| Mistake | Better practice |
|---|---|
| Assume one previous role is the “best” starting point. | Match your existing skills to the actual target responsibilities. |
| Learn every framework before building anything. | Learn the necessary primitives, then use one suitable implementation. |
| Treat sample thresholds as accepted industry standards. | Set criteria from task consequences, uncertainty and the available evidence. |
| Assume the model is always or never the bottleneck. | Trace the failure and compare plausible causes. |
| Equate a fixed version with permanent reproducibility. | Record artifacts and contracts, while planning for provider retirement or behavior changes. |
| Collect ratings without reviewing failures. | Connect feedback to traces and investigate representative cases. |
| Describe labor time saved as immediate payroll savings. | State whether time is redeployed, hiring is avoided or cash expense actually changes. |
| Publish employer data to make a portfolio realistic. | Use permitted public, synthetic or properly approved sanitized material. |
| Treat a salary range as a market average. | Check role, level, location, currency and whether the figure is base pay or total compensation. |
| Expect a certificate to guarantee an interview. | Pair learning with demonstrable implementation and clear explanation. |
Resources and final notes
The course guide connects each subject to current external study and original practice assignments. The hiring-evidence chapter explains the limits of job-posting and compensation data. The glossary provides concise definitions for revision.
- Choose a target scope you can explain and genuinely want to own.
- Build on your previous experience while being precise about the new skills you have demonstrated.
- Prepare one coherent piece of evidence with failures, measurements and tradeoffs before expanding the portfolio.
- Use feedback from mock interviews and real applications to revise the gap map.
- Keep claims about employment, responsibility, results and credentials truthful. Preparation improves readiness; hiring also depends on role fit, competition and the employer's decisions.