The idea looks simple: throw all your notes in a folder, let the agent loose in there, and get an assistant that remembers everything. In practice, the agent finds the wrong things, confuses documents, and forgets what was said five minutes ago. The reason is how its memory is structured.

The Agent Has No Memory

A model has only a context window, and it lives for one session. It contains messages and files from the current conversation, rules, skills, and short notes from past sessions, if the tool keeps them. Everything else doesn't exist for the model.

What is commonly called agent memory consists of two parts: external storage and a mechanism that retrieves the necessary pieces from it and puts them in the context. This mechanism is called RAG.

What Is RAG

RAG stands for Retrieval-Augmented Generation. There are three steps in the name:

  • Retrieval — find fragments in storage that match the question.
  • Augmented — add them to the query.
  • Generation — get the model's answer based on the added text.

This is how the context limitation is circumvented. You can't put an entire book or notes base in a query: either it won't fit or it will be expensive. RAG only gives the model an extract.

How It Works Step by Step

Data Preparation. Text is cut into small fragments, called chunks. Each chunk passes through an embedding model, which converts it into a long set of numbers, an embedding. The numbers describe the meaning of the text. Embeddings are stored in a vector database.

Query Processing. The user's question is also converted into an embedding. The database compares it with saved ones and finds the closest matches. This is semantic search: matching by meaning, not by words.

Context Assembly. The text of the found chunks is substituted into the query and sent to the model.

Any memory tool for an agent is built on these three steps. The details differ.

RAG Is Development, Not a Button

There's no ready-made "RAG for Claude Code" in one file. It's a separate engineering discipline, and there are two paths: build the pipeline yourself or connect a ready-made solution. Both are below.

Building It Ourselves: The Basics

It's most convenient to start with the framework LlamaIndex. It's a set of ready-made modules for each RAG task. It can handle the entire pipeline on its own, or it can tie together more powerful tools at individual stages. There are three stages: chunking, embedding, vector database.

Chunking

Search goes by chunks, so the slicing determines everything else. Main strategies:

  • By Size. Text is cut every 400–512 tokens regardless of meaning. Fast and predictable, but a sentence or table may end up cut. Overlap saves the day: 10–20% of the end of one chunk is repeated at the beginning of the next.
  • Recursive. Cuts along natural boundaries, first by paragraphs. You can set custom delimiters, so the method works for code too.
  • By Sentences. Accumulates whole sentences until it hits the limit. The thought is not interrupted, but long sentences create chunks that are too large. Good for dialogues and Q&A.
  • By Pages. One page — one chunk. Works for PDFs, reports, and presentations, where it's important to keep a table near its caption.
  • Semantic. Computes an embedding for each sentence and cuts where adjacent sentences diverge in meaning. More precise, but each embedding model call costs time and money.
  • By Model. The document is given to a language model, and it decides where the boundaries are, and also writes a brief description of each chunk. The most expensive and slowest method.
  • Late Chunking. The model first processes the document in full, so each word gets the context of the entire text, and only then the text is divided into chunks. This is an additional layer on top of regular slicing. You need an embedding model with a long context.

What to Choose. Start with recursive slicing at 400–512 tokens with 10–20% overlap. There's a breakdown showing that simple recursive slicing at 512 tokens outperforms complex methods in answer accuracy. Semantic and model-based chunking make sense for large databases for analytical tasks.

Usually two things give more benefit:

  • chunk metadata (tags, dates, document type) and filtering by it;
  • cleaning sources of garbage before slicing.

Tools. LlamaIndex is enough for a simple pipeline. If you want more strategies and speed, check out the Chonkie library. For PDFs and office files, you'll need:

  • pytesseract, pdf2image, PIL — basic toolkit for recognizing scans.
  • LlamaParse — a parser from the LlamaIndex ecosystem, preserves tables and multicolumn layouts.
  • Unstructured — extracts text from PDFs, Word, PowerPoint, HTML, and emails while preserving headers and lists.
  • Docling — a similar tool with a wide range of formats.
  • pdfplumber — returns text along with word coordinates, so you can show where in the document a quote came from.

LlamaIndex has a parser for Markdown notes that preserves the header structure in chunk metadata.

Embedding

Don't get stuck at this stage: there are many embedding models, but the difference between them for a personal base is small. A local multilingual model like paraphrase-multilingual-MiniLM-L12-v2 or nomic-embed-text will do.

A special case is multimodal models, like Gemini Embedding 2. They translate text, images, audio, video, and PDFs into a single semantic space, and you don't need to build a separate pipeline for each format. The model is paid.

Vector Database

This is where embeddings are stored and searched. Databases are local or cloud-based. Local keeps vectors in RAM, so you need to configure it to save to disk. Otherwise, after a restart, you'll have to recalculate everything. In the cloud, the provider takes care of this.

  • Pinecone — cloud-based, no administration required. Has hybrid search and metadata filters. Paid, not available locally.
  • Qdrant — open-source and fast, with advanced filtering. A good choice for local installation.
  • Chroma — lightweight, installed with one command. Good for prototypes and small personal databases.
  • Milvus — for very large volumes, works locally and in the cloud.
  • pgvector — PostgreSQL extension. Convenient if your data is already in Postgres.

From here on, you can build the pipeline with the agent: you already know the stages and terms and can describe what you want to get.

What Other Types of RAG Are There

Graph RAG

Regular RAG cuts text into independent chunks and loses connections between them. The graph version extracts entities and relationships between them from the text and stores them in a graph database. The agent sees not just a fragment but what it's connected to. The approach works well on coherent texts: books, articles, transcripts, personal notes. It's less suitable for scattered tables.

Important note for Obsidian users: links like [[note]] don't help the graph pipeline. To it, this is just regular text, it builds the connections itself.

Tools: Microsoft GraphRAG as a framework, Neo4j as a graph database, Gephi for visualization. For a deep dive, there's the book Essential GraphRAG.

Agentic RAG

It's a wrapper on top of any pipeline. The agent controls the search itself: it parses the query, selects a tool, evaluates what's found, and repeats the search with different wording if there's not enough data. The cycle continues until there's enough context for an answer.

Tools: LangChain for simple cases and LangGraph when you need loops, branching, and self-checking.

Hierarchical RAG

A document is divided into large parent fragments, and those into small child fragments. You search by small ones because they have less noise, and the models return the parent fragment in full. This solves an old dilemma: small chunks are precise but lacking context, large ones are the opposite.

In LlamaIndex, there's RaptorPack, as well as the pair HierarchicalNodeParser and AutoMergingRetriever: the first cuts the document into multiple levels, the second merges found small pieces into parent ones.

Ready-Made Solutions

Almost all the tools mentioned can be connected to the agent via MCP.

  • Mem0 — turnkey pipeline, cloud or local. Combines vector search, graph, and metadata, can rerank results. Memory is populated automatically: the system extracts key facts and solutions from sessions. The graph part is noticeably simpler than the vector part.
  • QMD — local search across Markdown notes. A local model expands the query, then word-based and semantic search run in parallel, results are combined and reranked. Chunking respects Markdown structure. Demands a lot of RAM.
  • OpenViking — stores memory as a file system with addresses like viking://. Each record has three levels of detail: a short summary, an overview, and the full text. Search first selects the appropriate directory, then goes deeper, making it easier to verify.
  • Graphiti — a wrapper around graph databases that ties facts to time: when an entity appeared, how relationships changed, when a fact became outdated. Convenient for data that constantly changes, for example project history.
  • Hindsight — graph memory with four types of records: facts about the world, agent experience, opinions with confidence scores, and observations. Searches by meaning, words, graph, and time. Works right after installation, connects to Claude Code as a plugin.
  • Cognee — customizable knowledge engine: vectors, graph, and metadata. Accepts dozens of formats, builds entity graph on its own, spins up with a few commands.

Getting Started

  1. Build a minimal pipeline: recursive chunking, a local embedding model, Chroma or Qdrant.
  2. Test it on ten of your real questions.
  3. Add metadata and filters.
  4. Complicate things only where you spot a concrete search error.

Most of the tools in the article are free with local installation.