Learnastra AI SYSTEM DESIGNAnup Rai

Concept · Understand the mechanism

Memory Architectures

By Anup Rai7 min readReviewed September 2026

An AI application's memory architecture defines what information it retains, how it changes and how the application selects it for later use. Some memory is conversation-specific; some persists across sessions. The model only uses the information made available through its input, tools or learned parameters—it does not automatically read every record the application stores.

Start an interview with three separate questions: What is stored? Who may use it? When is it valid? A vector database answers none of these on its own.

Separate purpose, scope and implementation

Working, episodic, semantic and procedural memory are useful cognitive analogies in agent design. They are not a mandatory CPU-like L1–L4 hierarchy. Nor does one category require a particular database or have an intrinsic latency. Memory terminology and scope.

Category Meaning in this guide Example Possible representation
Working context Information selected for the current model interaction Current task, recent messages, retrieved evidence Messages and structured fields
Episodic memory Records of particular events or experiences A previous troubleshooting attempt and its outcome Event rows, documents, artifact references
Semantic memory Retained facts or assertions about entities A user's preferred explanation language Relational fields, documents or graph assertions
Procedural memory Retained instructions or methods A reviewed troubleshooting procedure Versioned instructions, code or selected examples

These categories can overlap. One completed exercise is an episode; a tested lesson derived from many exercises might inform a procedure. This promotion is an application decision, not an automatic consequence of storing an embedding.

Semantic memory does not mean semantic search. The former describes retained knowledge; the latter describes retrieval by meaning. A SQL lookup can retrieve a semantic-memory fact. Vector search can retrieve an episodic record.

Separate mechanism What it does What it does not establish
KV cache Reuses intermediate attention computations Durable user preferences or authoritative truth
Conversation checkpoint Preserves selected execution state The outcome of an unrecorded external write
Document retrieval Supplies evidence from a corpus Permission to treat its text as instructions
Model weights Encode learned statistical behavior An editable, individually deletable user-memory table
Business system of record Owns a domain's authoritative state That every cached copy is current

Begin with a small, concrete product

Consider an interview-practice tutor that should remember a learner's chosen language and completed exercises. This is a design exercise, not a claim about a deployed product.

Functional requirements

  1. Resume an unfinished practice session.
  2. Reuse explicit preferences in later sessions.
  3. Recall relevant prior attempts and feedback.
  4. Let the learner inspect, correct and remove remembered information.
  5. Distinguish verified exercise results from model-generated interpretations.

Non-functional requirements

  1. Enforce account and organization boundaries on every memory operation.
  2. Keep response latency and memory-read cost within the agreed budget.
  3. Preserve source, time, version and correction history where required.
  4. Propagate deletions and corrections to derived retrieval indexes and caches.
  5. Avoid silently converting temporary session choices into permanent preferences.

Start with a relational database: a preferences table, exercise-attempt records and session state. Load preferences by authenticated account ID and retrieve recent attempts by exercise/topic. This may satisfy the product without a vector index or graph.

The first failure might be a query such as “Which earlier problem had the same failure pattern?” Exact topic labels may miss a relevant exercise under a different name. Add semantic retrieval over attempt summaries if evaluation shows a useful gain. Add relationship traversal when the queries actually depend on linked concepts, prerequisites or projects.

Design the write path and the read path separately

Architecture / visual model
flowchart LR A[Authenticated interaction or event] --> P[Apply scope and retention policy] P --> E[Extract candidate assertions if needed] E --> V[Validate source, meaning and version] V --> S[(Authoritative records with provenance)] S --> I[Update derived indexes] Q[New scoped task] --> R[Retrieve permitted and current records] S --> R I --> R R --> B[Select evidence within context budget] B --> M[Model interaction] C[Correction or deletion] --> S C --> I
Read diagram source
flowchart LR
    A[Authenticated interaction or event] --> P[Apply scope and retention policy]
    P --> E[Extract candidate assertions if needed]
    E --> V[Validate source, meaning and version]
    V --> S[(Authoritative records with provenance)]
    S --> I[Update derived indexes]
    Q[New scoped task] --> R[Retrieve permitted and current records]
    S --> R
    I --> R
    R --> B[Select evidence within context budget]
    B --> M[Model interaction]
    C[Correction or deletion] --> S
    C --> I

An explicit preference can be written through ordinary validated form/API code. A transcript may require an extraction model, but its output is a candidate assertion. “Use Java for this exercise” must not become “always prefers Java.” Attach the scope and source that justify the claim.

For reads, derive identity and allowed scopes from the authenticated application session. Never accept the model's proposed user_id as sufficient authorization. Apply access constraints before evidence reaches the model and revalidate returned records under the storage system's consistency contract.

Decide when writes become visible

Pattern Benefit Cost or failure to handle
Write on the request path Immediate acknowledgment and easier read-after-write behavior Adds latency; write failure affects the request
Extract/index asynchronously Keeps expensive processing off the response path Temporary stale retrieval, retries and backlog
Store authoritative change synchronously, index later Fast durable correction with cheaper search maintenance Reader must account for index lag

For the tutor, commit a language preference before confirming it to the learner. Indexing a long exercise transcript can happen later. If the next turn needs that transcript immediately, use the authoritative session record rather than waiting for search indexing.

Consolidation derives a more compact or useful representation from existing information. It can combine duplicate assertions or summarize episodes, but must retain enough provenance to explain and correct the result. Frequently retrieved information is not necessarily more truthful. An attacker can repeat a false claim; repeated retrieval can also reinforce the system's own mistake.

Resolve conflicts without inventing a truth hierarchy

Conflict Appropriate question Example response
Old and new explicit preference Do both apply to the same scope and period? Supersede the old global choice, retain a historical record if permitted
Profile versus temporary request Is this a session exception? Use the requested language for this exercise only
Generated summary versus scored result Which source owns this fact? Use the authoritative assessment record
Two uncertain extracted assertions Is either sufficiently supported? Retain uncertainty or ask for clarification

Semantic memory is neither immutable nor inherently authoritative. A job, address or preference can change. Record when a fact applies and when the system learned it. “Newest timestamp wins” is insufficient if the new item repeats an old document or applies to a different project.

Estimate the footprint

Illustrative assumptions: 100,000 learners, 40 retained records per learner, 600 bytes of text/metadata per record, and one 768-dimensional float32 embedding per record.

Quantity Calculation Raw size
Records 100,000 × 40 4 million
Text and metadata 4 million × 600 bytes 2.4 GB
Embeddings 4 million × 768 × 4 bytes 12.288 GB
Combined, three copies (2.4 + 12.288) × 3 44.064 GB

These decimal GB figures exclude indexes, database overhead, logs and backups. Do not embed fields that only need exact lookup. Evaluate whether embeddings, raw transcripts and replicas have the same retention requirements.

Latency comes from the concrete queries, index, network and load. There is no universal rule that semantic memory takes over 500 ms or episodic memory takes 100–300 ms. Measure the path your design uses.

Test memory as a lifecycle

  1. Extraction: Did the stored assertion preserve negation, scope and uncertainty?
  2. Update: Did a correction become effective without losing unrelated facts?
  3. Retrieval: Did the system find the right records and exclude forbidden ones?
  4. Use: Did those records improve the task outcome without irrelevant personalization?
  5. Deletion: Did the information disappear from active records and derived views under the defined retention policy?
  6. Recovery: Can indexing retries or restored backups resurrect a removed assertion?

An architecture that retrieves many records can still be worse if those records are stale or distracting. Compare with a no-memory baseline and a simple structured-profile baseline.

Interview practice

Q1: Is a three-tier memory architecture the industry standard?

No. Cognitive categories help describe information, but scope, authority, storage and retrieval are separate choices. Explain the required behaviors and then choose the smallest architecture that supports them.

Q2: Why not put the whole user history in the prompt?

It may exceed the model or application budget, increase processing cost and introduce irrelevant or obsolete evidence. Full history can still be a valid baseline for small workloads. Compare it with selective retrieval and evaluate task accuracy as well as cost.

Q3: Must semantic memory use a graph database?

No. A preference can be a versioned relational row. A graph becomes useful when relationships and traversal are central to the required queries; it adds identity-resolution and maintenance costs.

Q4: Where should an exercise score live?

In the assessment system's authoritative record. Memory can retain a reference or derived learning summary, but a model's recollection must not silently replace the scored result.

Q5: How would you prevent cross-account recall?

Derive scope from authenticated identity, enforce access in the read/write path, and test caches, indexes, exports and administrative operations as well as the main database. A namespace field without enforced checks is only a label.

Q6: What makes a good memory-service abstraction?

Explicit contracts for source, scope, version, freshness, correction and deletion, plus observable latency and cost. A generic remember(text) method hides too much if the product requires reliable updates or sensitive isolation.

Final notes

Recall card: Store deliberately → qualify the assertion → enforce scope → retrieve selectively → correct and forget reliably.

The Generative Agents research illustrates an architecture combining a memory stream, retrieval and reflection. It is a research design, not proof that every production product needs that arrangement.

Next: Short-term context, then long-term memory.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Loop Engineering
NEXT LESSONShort-Term Context Management →

Explore the diagram