Skip to content
Sponsor

RAG with Vectors

Retrieval-augmented generation in two branches of one graph. The first splits a notes file into chunks, embeds each one and stores it. The second embeds your question, finds the closest chunks, and hands them to the model with the question.

Download the graph and save it in user/graphs/, or build it below. The graph is MIT-0: copy it and license what you build from it however you like.

Terminal window
ollama pull bge-m3
ollama pull gemma4:e4b

Put your notes in a text or Markdown file at user/data/files/notes.md. That folder is the only one the file nodes read.

The downloaded graph picks no model, because a model can cost money and Boltjar never picks one for you. It stays Off until both Embed nodes have one, and its problems list says pick a model for Embed for each. Pick bge-m3 (Ollama) on Embed chunk and on Embed question, and a chat model such as Gemma 4 e4b (Ollama) on the LLM; without that one, the LLM answers with a mock reply.

Node Settings
Manual, named Index none
Read File, named Notes path notes.md
Chunk, named Chunks size 800, overlap 120, by chars
Vector Store none
Vectors, named Clear operation clear, namespace notes
For-each, named Each chunk none
Embed, named Embed chunk model: pick bge-m3 (Ollama)
Vectors, named Add operation index, namespace notes
From To
Index trigger Clear trigger
Vector Store vectors Clear vectors
Clear trigger Each chunk trigger
Notes text Chunks text
Chunks out Each chunk list
Each chunk each Embed chunk trigger
Each chunk item Embed chunk text
Embed chunk trigger Add trigger
Vector Store vectors Add vectors
Embed chunk embedding Add embedding
Each chunk item Add text
Add trigger Each chunk loop

Read it as a loop: Clear empties the notes namespace, then For-each hands out the chunks one at a time. Each chunk is embedded and added, and Add’s trigger goes back into loop to release the next chunk. Clearing first means a second run replaces the notes instead of adding them twice.

Clear takes no embedding: a Vectors node shows its embedding input only for search and index.

Node Settings
Chat Input, named Question none
Embed, named Embed question model: pick bge-m3 (Ollama)
Vectors, named Relevant notes operation search, namespace notes, top k 4, min score 0.3
Template, named Prompt the text below
LLM model: pick Gemma 4 e4b (Ollama), or another model you pulled
Chat none
Answer the question using the notes below. If they do not cover it, say so.
Notes:
{relevant_notes}
Question: {question}
From To
Question trigger Embed question trigger
Question text Embed question text
Embed question trigger Relevant notes trigger
Vector Store vectors Relevant notes vectors
Embed question embedding Relevant notes embedding
Relevant notes trigger Prompt trigger
Relevant notes results Prompt relevant_notes
Question text Prompt question
Prompt trigger LLM trigger
Prompt out LLM prompt

Plus the four Chat wires from the first graph: Question into user and user_trigger, the LLM into reply and reply_trigger.

The Prompt’s tags are named after the nodes their wires come from: Relevant notes results becomes {relevant_notes}, Question text becomes {question}. The question fires each node once down the line, through the Prompt to the LLM, so one question runs the model once.

Once the models are picked, press On. The Manual trigger fires as the graph turns On, so indexing starts right away; press Fire on the Index node to run it again after you edit notes.md. Then ask a question in the Question node.

Each hit reaches the prompt with its text and its score. If the answer says the notes do not cover something they do, check that Ollama was running when the Index node fired: when Ollama does not answer, Embed gives an empty vector and the search finds nothing, so press Fire on the Index node again once it runs. A bge-m3 that is not pulled keeps the graph Off instead, and its problems list says to pull the model in Settings › AI Providers.

  • More or fewer hits. Raise top k for broader answers, or min score to keep only close matches.
  • Many files. Put a List Dir and a second For-each in front of Read File to index a folder.
  • A sharper order. Put a Rerank between Relevant notes and the Prompt. It needs a rerank model picked on it and a rerank service running over HTTP; Boltjar does not bundle one.
  • Remember the source. Set Add’s metadata to ref = notes.md, and delete with that ref removes the file’s chunks later.