KevsRobots Learning Platform
30% Percent Complete
By Kevin McAleer, 6 Minutes
Page last updated August 23, 2026

An Obsidian vault is just a folder of markdown files. No database, no proprietary format - which is the whole reason Obsidian is worth building tools for.
That does mean a few things are lurking in there that you do not want in your index.
| Folder | Why skip it |
|---|---|
.obsidian |
Plugin config, themes, workspace layout - thousands of lines of JSON and the odd stray markdown file |
.trash |
Notes you deleted. Retrieving one would be genuinely confusing |
.git |
If you version control your vault |
.smart-env |
Cache folder left by some AI plugins |
templates |
Placeholder text like {{title}} that matches everything and means nothing |
Those are already in SKIP_DIRS in config.py. Add your own as you find them.
Everything downstream works with Note objects rather than raw file paths, so the parsing happens exactly once:
# vault_rag/vault.py
"""Read an Obsidian vault from disk and turn it into Note objects."""
import hashlib
from dataclasses import dataclass, field
from pathlib import Path
from .config import SKIP_DIRS
@dataclass
class Note:
"""One markdown file from the vault, already parsed."""
path: Path
relative_path: str
title: str
body: str
frontmatter: dict = field(default_factory=dict)
tags: list[str] = field(default_factory=list)
links: list[str] = field(default_factory=list)
mtime: float = 0.0
@property
def folder(self) -> str:
parent = Path(self.relative_path).parent
return "" if str(parent) == "." else str(parent)
@property
def content_hash(self) -> str:
"""Fingerprint of the note's text - changes when the note changes."""
return hashlib.sha256(self.body.encode("utf-8")).hexdigest()[:16]
frontmatter, tags and links are declared but stay empty until the next lesson. The two properties are worth a closer look.
folder turns robots/smars.md into robots, and a note at the vault root into "". Obsidian users organise by folder constantly, so this makes a great filter.
content_hash is the key to fast re-indexing. Hash the body, keep the first 16 hex characters, and you have a fingerprint that changes whenever the noteβs text changes. In lesson 9 we compare that against what is stored in the index and skip anything unchanged. Note that it hashes the body, not the file - so touching a file without editing it does not trigger a re-index.
# vault_rag/vault.py (add below the Note class)
def load_note(path: Path, vault_path: Path) -> Note:
"""Read and parse one markdown file."""
text = path.read_text(encoding="utf-8", errors="replace")
relative_path = str(path.relative_to(vault_path))
return Note(
path=path,
relative_path=relative_path,
title=path.stem,
body=text.strip(),
mtime=path.stat().st_mtime,
)
Two small details doing real work:
errors="replace" stops one badly encoded file from killing the whole indexing run. A note pasted from a Windows machine years ago can contain bytes that are not valid UTF-8; replace swaps them for a placeholder character and carries on.
relative_path is what we store in the index, not the absolute path. Move your vault to a different machine and the index still makes sense.
# vault_rag/vault.py (add below load_note)
def iter_notes(vault_path: Path):
"""Yield every indexable note in the vault, in a stable order."""
vault_path = Path(vault_path).expanduser()
if not vault_path.is_dir():
raise FileNotFoundError(f"Vault not found: {vault_path}")
for path in sorted(vault_path.rglob("*.md")):
parts = set(path.relative_to(vault_path).parts)
if parts & SKIP_DIRS:
continue
note = load_note(path, vault_path)
if note.body:
yield note
Four decisions in nine lines:
It is a generator. yield, not return [...]. A big vault has thousands of notes and there is no reason to hold them all in memory at once - the indexer processes them one at a time and lets each go.
sorted() makes the order deterministic. Without it, rglob returns whatever order the filesystem feels like, which differs between macOS and Linux and makes debugging miserable.
The skip check uses a set intersection. parts is every path component - ("robots", "old", "smars.md") - so parts & SKIP_DIRS catches a skipped folder at any depth, not just the top level. A templates folder nested three levels down still gets skipped.
Empty notes are dropped. An empty file produces an empty embedding and pollutes results. Obsidian creates these constantly when you click βnew noteβ and then wander off.
Point it at your real vault:
# try_it.py - a scratch script in the project root, next to vault_rag/
from vault_rag.config import VAULT_PATH
from vault_rag.vault import iter_notes
notes = list(iter_notes(VAULT_PATH))
print(f"{len(notes)} notes")
# The five biggest - these are the ones chunking has to handle well
for note in sorted(notes, key=lambda n: len(n.body), reverse=True)[:5]:
print(f"{len(note.body):7,} chars {note.relative_path}")
Then try these:
len(notes) against find ~/YourVault -name "*.md" | wc -l. The difference is your skipped and empty notes. Is it the number you expected?MIN_CHUNK_SIZE."attachments" to SKIP_DIRS and see whether your count changes.content_hash twice, then edit the note and print it again. Watch it change.FileNotFoundError: Vault not found: ~/Obsidian/MyVaultiter_notes calls .expanduser() for you, so this means the path genuinely does not exist - check for a typo or a space in the folder name.Why: Path("~/x") is a literal path with a ~ directory in it, which almost certainly does not exist.
.obsidian still appear.path.relative_to(vault_path).parts and not path.parts.Why: path.parts includes every component of the absolute path. If your vault happens to live under a folder called templates, you would skip the entire vault.
attachments folder full of PDFs or images. rglob("*.md") still has to stat everything it walks past.Why: rglob walks the entire directory tree. Adding heavy asset folders to SKIP_DIRS does not help here because the walk happens first - if it is a real problem, use os.walk and prune directories as you go.
UnicodeDecodeError on one specific file.errors="replace" from read_text.
You can use the arrows β β on your keyboard to navigate between lessons.
Comments