RAG with Vectors
Retrieval-augmented generation in two branches of one graph. The first splits a notes file into chunks, embeds each one and stores it. The second embeds your question, finds the closest chunks, and hands them to the model with the question.
Download the graph and save it in user/graphs/, or build it below. The graph is MIT-0: copy it and license what you build from it however you like.
1. Get ready
Section titled “1. Get ready”ollama pull bge-m3ollama pull gemma4:e4bPut your notes in a text or Markdown file at user/data/files/notes.md. That folder is the only one the file nodes read.
The downloaded graph picks no model, because a model can cost money and Boltjar never picks one for you. It stays Off until both Embed nodes have one, and its problems list says pick a model for Embed for each. Pick bge-m3 (Ollama) on Embed chunk and on Embed question, and a chat model such as Gemma 4 e4b (Ollama) on the LLM; without that one, the LLM answers with a mock reply.
2. The indexing branch
Section titled “2. The indexing branch”| Node | Settings |
|---|---|
| Manual, named Index | none |
| Read File, named Notes | path notes.md |
| Chunk, named Chunks | size 800, overlap 120, by chars |
| Vector Store | none |
| Vectors, named Clear | operation clear, namespace notes |
| For-each, named Each chunk | none |
| Embed, named Embed chunk | model: pick bge-m3 (Ollama) |
| Vectors, named Add | operation index, namespace notes |
| From | To |
|---|---|
Index trigger |
Clear trigger |
Vector Store vectors |
Clear vectors |
Clear trigger |
Each chunk trigger |
Notes text |
Chunks text |
Chunks out |
Each chunk list |
Each chunk each |
Embed chunk trigger |
Each chunk item |
Embed chunk text |
Embed chunk trigger |
Add trigger |
Vector Store vectors |
Add vectors |
Embed chunk embedding |
Add embedding |
Each chunk item |
Add text |
Add trigger |
Each chunk loop |
Read it as a loop: Clear empties the notes namespace, then For-each hands out the chunks one at a time. Each chunk is embedded and added, and Add’s trigger goes back into loop to release the next chunk. Clearing first means a second run replaces the notes instead of adding them twice.
Clear takes no embedding: a Vectors node shows its embedding input only for search and index.
3. The question branch
Section titled “3. The question branch”| Node | Settings |
|---|---|
| Chat Input, named Question | none |
| Embed, named Embed question | model: pick bge-m3 (Ollama) |
| Vectors, named Relevant notes | operation search, namespace notes, top k 4, min score 0.3 |
| Template, named Prompt | the text below |
| LLM | model: pick Gemma 4 e4b (Ollama), or another model you pulled |
| Chat | none |
Answer the question using the notes below. If they do not cover it, say so.
Notes:{relevant_notes}
Question: {question}| From | To |
|---|---|
Question trigger |
Embed question trigger |
Question text |
Embed question text |
Embed question trigger |
Relevant notes trigger |
Vector Store vectors |
Relevant notes vectors |
Embed question embedding |
Relevant notes embedding |
Relevant notes trigger |
Prompt trigger |
Relevant notes results |
Prompt relevant_notes |
Question text |
Prompt question |
Prompt trigger |
LLM trigger |
Prompt out |
LLM prompt |
Plus the four Chat wires from the first graph: Question into user and user_trigger, the LLM into reply and reply_trigger.
The Prompt’s tags are named after the nodes their wires come from: Relevant notes results becomes {relevant_notes}, Question text becomes {question}. The question fires each node once down the line, through the Prompt to the LLM, so one question runs the model once.
4. Run it
Section titled “4. Run it”Once the models are picked, press On. The Manual trigger fires as the graph turns On, so indexing starts right away; press Fire on the Index node to run it again after you edit notes.md. Then ask a question in the Question node.
Each hit reaches the prompt with its text and its score. If the answer says the notes do not cover something they do, check that Ollama was running when the Index node fired: when Ollama does not answer, Embed gives an empty vector and the search finds nothing, so press Fire on the Index node again once it runs. A bge-m3 that is not pulled keeps the graph Off instead, and its problems list says to pull the model in Settings › AI Providers.
Change it
Section titled “Change it”- More or fewer hits. Raise
top kfor broader answers, ormin scoreto keep only close matches. - Many files. Put a List Dir and a second For-each in front of Read File to index a folder.
- A sharper order. Put a Rerank between Relevant notes and the Prompt. It needs a rerank model picked on it and a rerank service running over HTTP; Boltjar does not bundle one.
- Remember the source. Set Add’s
metadatatoref = notes.md, anddeletewith thatrefremoves the file’s chunks later.