Context
Running a local model is only part of the problem. To work with private knowledge, the system needs to locate relevant information and provide it to the model when required.
This experiment explores the construction of a fully local RAG architecture: from document ingestion and chunking to embeddings, vector search and context retrieval for a local LLM.
Goal
Build a small and understandable foundation before adding more layers.
The architecture being explored deliberately starts with a simple pipeline:
Documents
↓
Parsing
↓
Chunking
↓
Embeddings
↓
Vector store
↓
Retrieval
↓
Context
↓
Local LLM
The priority is not to hide the process behind a framework, but to understand what happens at each stage and which decisions actually affect retrieval quality.
What we are investigating
- document chunking strategies
- local embedding generation
- vector storage and search
- relevant context selection
- the relationship between retrieval and context windows
- integration with local LLM inference
Principle
Knowledge and inference are separate problems.
The model does not need to contain all the information we use. Part of the knowledge can remain outside the model, stored persistently and retrieved only when it becomes relevant.
This makes it possible to treat knowledge as a layer independent from the inference engine.
Status
Active.
The foundation is still under construction, and the next phase is to refine the RAG pipeline before using it as the basis for more specific systems.