KevsRobots Learning Platform
25% Percent Complete
By Kevin McAleer, 7 Minutes
Page last updated August 23, 2026

Before we point ChromaDB at a real vault, letβs spend a lesson just playing with it. It has a small API and you can learn nearly all of it in twenty minutes.
ChromaDB stores documents. A document is:
Documents live in a collection, which is a bit like a table. We will use exactly one collection for the whole vault.
That is genuinely it.
# scratch_chroma.py - a throwaway script for exploring the API
import chromadb
client = chromadb.PersistentClient(path="./scratch_db")
collection = client.get_or_create_collection(
name="notes",
configuration={"hnsw": {"space": "cosine"}},
)
print(collection.count())
Three things worth pausing on:
PersistentClient writes to disk, so the index survives between runs. There is also chromadb.EphemeralClient() which keeps everything in memory - great for tests, useless for our tool.
get_or_create_collection does what it says, which makes it safe to call on every run. create_collection raises if the collection already exists.
configuration={"hnsw": {"space": "cosine"}} sets the distance metric. Chromaβs default is squared L2, which works but produces distances that are harder to reason about. Cosine is the right choice for text, and you must set it when the collection is created - you cannot change it later without rebuilding.
# scratch_chroma.py (continued)
collection.upsert(
ids=["note1-0", "note2-0", "note3-0"],
documents=[
"The Pico W has a CYW43439 wireless chip and runs MicroPython.",
"My SMARS robot uses two N20 motors and a DRV8833 driver.",
"Sourdough starter needs feeding every 12 hours in warm weather.",
],
metadatas=[
{"note": "Pico.md", "folder": "electronics"},
{"note": "SMARS.md", "folder": "robots"},
{"note": "Bread.md", "folder": "kitchen"},
],
)
print(collection.count()) # 3
Notice we never mentioned embeddings. Chroma spotted that we passed documents rather than embeddings and ran the default model over them for us.
Chroma has an add() method too. Use upsert() instead, always.
The difference bites in a way you will not notice until it has cost you an hour:
# scratch_chroma.py (continued)
collection.add(ids=["note1-0"], documents=["Completely different text"])
print(collection.get(ids=["note1-0"], include=["documents"])["documents"])
# ['The Pico W has a CYW43439 wireless chip and runs MicroPython.']
add() saw an id it already had and silently kept the old document. No exception, no warning in your output. Edit a note, re-index with add(), and your index quietly still holds the old version.
upsert() overwrites. That is what you want every single time.
# scratch_chroma.py (continued)
results = collection.query(
query_texts=["which board has wifi?"],
n_results=2,
include=["documents", "metadatas", "distances"],
)
print(results["ids"])
print(results["distances"])
[['note1-0', 'note2-0']]
[[0.5076934695243835, 0.8984732627868652]]
The Pico note wins, even though the question said βwifiβ and the note said βwireless chipβ.
Those double brackets trip up everyone. query_texts is a list - you can ask several questions in one call - so every result is a list-of-lists, one inner list per query.
We only ever ask one thing at a time, so you will see [0] all over our code:
# scratch_chroma.py (continued)
for document, metadata, distance in zip(
results["documents"][0],
results["metadatas"][0],
results["distances"][0],
):
print(f"{distance:.3f} {metadata['note']} {document[:40]}...")
include=["documents", "metadatas", "distances"] is not optional politeness - it is a real optimisation. Leave it off and older Chroma versions hand back the full 384 number embedding for every hit, which you then throw away.
This is where Chroma earns its keep. You can constrain the search before the similarity comparison happens:
# scratch_chroma.py (continued)
results = collection.query(
query_texts=["motors"],
n_results=5,
where={"folder": "robots"},
)
print(results["ids"]) # [['note2-0']]
The operators available on metadata are:
| Operator | Meaning | Example |
|---|---|---|
$eq / $ne |
Equals, not equals | {"folder": {"$eq": "robots"}} |
$gt $gte $lt $lte |
Numeric comparison | {"mtime": {"$gte": 1700000000}} |
$in / $nin |
Value is (not) in a list | {"folder": {"$in": ["robots", "electronics"]}} |
$contains / $not_contains |
List field does (not) contain a value | {"tags": {"$contains": "pico"}} |
$and / $or |
Combine clauses | {"$and": [clause1, clause2]} |
A bare {"folder": "robots"} is shorthand for {"folder": {"$eq": "robots"}}.
$contains is for list values, not substrings. If tags is ["pico", "micropython"] then {"tags": {"$contains": "pico"}} matches. If folder is the string "electronics" then {"folder": {"$contains": "electr"}} matches nothing - there is no substring operator for metadata.
Separately from metadata, you can filter on the chunk text itself:
# scratch_chroma.py (continued)
collection.query(
query_texts=["anything"],
n_results=5,
where_document={"$contains": "MicroPython"},
)
where_document supports $contains, $not_contains and $regex. Here $contains is a substring match, and it is case sensitive - "micropython" finds nothing. Use $regex with an inline flag when you need to ignore case:
where_document={"$regex": "(?i)micropython"}
We will use this in lesson 12 to rescue searches for exact part numbers, which is the thing embeddings are worst at.
get() is the non-semantic sibling of query() - no question, no ranking, just βgive me the rows matching this filterβ:
# scratch_chroma.py (continued)
collection.get(where={"note": "Pico.md"}, include=["metadatas"])
collection.get(ids=["note1-0"])
collection.get() # everything
delete() takes the same arguments:
# scratch_chroma.py (continued)
collection.delete(where={"note": "Bread.md"})
collection.delete(ids=["note1-0"])
Deleting by where is exactly how we will handle an edited note in lesson 9 - drop all of that noteβs chunks, then write the new ones.
n_results=100 on a collection holding 3 documents. Chroma returns 3, not an error - handy, because it means you never have to clamp n_results yourself.{"tags": ["a", "b"]} and filter with {"tags": {"$contains": "a"}}. Then try {"tags": {"$in": ["a"]}} and watch it return nothing. That distinction matters in lesson 11.scratch_db folder and re-run your script. Everything rebuilds from nothing - the database really is just that directory.Expected a name containing 3-512 characters from [a-zA-Z0-9._-], starting and ending with a character in [a-zA-Z0-9]"v" fails; "vault_notes" is fine.Why: Chroma validates collection names because they become directory names on disk.
Expected metadata to be a non-empty dict, got 0 metadata attributes{} as a metadata entry. Pass at least one key, or leave the whole metadatas argument off.Why: An empty dict is almost always a bug in your code rather than a deliberate choice, so Chroma refuses it.
None for a document you definitely gave metadata to.None value in the dict - something like {"title": None} when a note had no title.Why: A None value causes the entire metadata dictionary to be stored as None. Coerce your values to strings, or drop the key, before writing. We handle this properly in lesson 8.
configuration={"hnsw": {"space": "cosine"}} when the collection was created.
You can use the arrows β β on your keyboard to navigate between lessons.
Comments