Retrieval and RAG
Build knowledge bases, ingest documents, retrieve context, and expose retrieval as a Namzu tool using the public @namzu/sdk RAG surface.
The SDK ships a complete retrieval path: chunk content, embed it, store vectors, retrieve relevant chunks, assemble context, and optionally expose the result as a tool the model can call.
1. The RAG Pipeline
The public exports line up as one pipeline:
| Stage | Owns | Main exports |
|---|---|---|
| Chunking | split documents into searchable units | TextChunker, DEFAULT_CHUNKING_CONFIG |
| Embeddings | convert text into vectors | EmbeddingProvider, OpenRouterEmbeddingProvider |
| Vector store | persist searchable vectors | InMemoryVectorStore, VectorStore |
| Retrieval | rank relevant chunks | DefaultRetriever, DEFAULT_RETRIEVAL_CONFIG |
| Knowledge base | one scoped corpus with ingest/query methods | DefaultKnowledgeBase |
| Context assembly | turn search hits into prompt-safe text | assembleRAGContext, DEFAULT_RAG_CONTEXT_CONFIG |
| Tool adapter | expose retrieval to the model | createRAGTool() |
2. Runnable Local Example
This example avoids external services by using a tiny in-memory embedding provider that satisfies the public EmbeddingProvider contract:
3. Expose the Knowledge Base as a Tool
Once the knowledge base exists, adapt it into a standard tool definition:
Important public behavior:
- the tool name is
knowledge_search - tool input uses snake_case fields such as
knowledge_base_idandtop_k - the tool returns assembled context text in
output - source metadata is returned in
data.sources
4. Choose Retrieval Mode Intentionally
DefaultRetriever supports three public modes:
| Mode | Best for | Tradeoff |
|---|---|---|
vector | semantic similarity | depends entirely on embedding quality |
keyword | term-heavy exact matching | weaker semantic recall |
hybrid | mixed semantic plus lexical retrieval | more work, but the safest default for many docs corpora |
You can also pass threadMessages in RetrievalQuery. The retriever expands the query with recent thread context before search, which helps when the user asks short follow-up questions.
5. Chunking Strategy Changes the Whole System
The chunking surface is not cosmetic. It shapes what retrieval can find.
| Strategy | Good default for |
|---|---|
fixed | uniform chunks and low-complexity ingestion |
sentence | short factual corpora |
paragraph | documentation and prose-heavy material |
recursive | mixed content where you want progressively smaller splits |
For documentation corpora, paragraph or recursive is usually the best starting point because they preserve more semantic shape than fixed slices.
6. Knowledge Base Scope Matters
DefaultKnowledgeBase is tenant-scoped. That is not just metadata. The underlying retriever and vector store filter on tenant identity, so one tenant's data does not bleed into another tenant's search results.
Use one knowledge base when:
- one corpus has one retention and retrieval policy
- one tenant owns the documents
- one tool should search one consistent namespace
Use multiple knowledge bases when:
- you want distinct corpora such as product docs vs. customer data
- you need different chunking or retrieval policies
- you want the model to choose a corpus explicitly through
knowledge_base_id
7. assembleRAGContext() Is Useful Even Without the Tool
If you want retrieval for a UI or an internal runtime layer rather than a tool call, you can stop at assembleRAGContext():
This is useful when:
- a server route wants to build prompt context manually
- you want to inspect sources before tool exposure
- retrieval is part of a larger orchestration path
8. Common Mistakes
| Mistake | Why it hurts |
|---|---|
expecting createRAGTool() to ingest documents for you | ingestion is owned by the knowledge base, not the tool wrapper |
keying the knowledgeBases map with the wrong ID | the tool resolves by knowledge-base ID, not arbitrary labels |
| forgetting tenant boundaries | retrieval results are scoped by tenant, so mixed-tenant data will not behave as one corpus |
| using camelCase tool input names | the tool schema uses knowledge_base_id and top_k |
assuming InMemoryVectorStore is a production persistence layer | it is great for tests, demos, and ephemeral workers, not long-lived storage |