A retrieval-augmented generation system built from the ground up to understand and experiment with the core components of RAG.
This project focuses on understanding and implementing RAG fundamentals before introducing higher-level frameworks or managed vector databases.
Current pipeline:
Documents
↓
Text Extraction
↓
Structure-Aware Chunking
↓
Embeddings
↓
Semantic Retrieval
↓
Retrieval Evaluation
The initial document used for development is a resume, but the ingestion and retrieval pipeline is being designed to support other document sources.
- PDF text extraction
- Text cleaning
- Structure-aware chunking
- Chunk metadata
- Persistent chunk storage
- Sentence Transformer embeddings
- Persistent embeddings
- Cosine-similarity retrieval
- Top-k semantic retrieval
- Initial retrieval evaluation
- Retrieval improvements
- Metadata filtering
- Hybrid search
- Reranking
- Vector database
- Context construction
- LLM generation
- Answer citations
- RAG evaluation
- FastAPI service
- Docker deployment
- CI/CD
The current retriever uses:
all-MiniLM-L6-v2- 384-dimensional embeddings
- normalized embeddings
- cosine similarity
- top-k retrieval
Initial evaluation:
| Metric | Score |
|---|---|
| Recall@3 | 0.750 |
| Precision@3 | 0.333 |
| MRR | 0.583 |
These results represent the initial baseline and will be used to measure improvements to the retrieval pipeline.
┌──────────────────┐
│ Documents │
└────────┬─────────┘
↓
┌──────────────────┐
│ Ingestion │
└────────┬─────────┘
↓
┌──────────────────┐
│ Chunking │
│ + Metadata │
└────────┬─────────┘
↓
┌──────────────────┐
│ Embeddings │
└────────┬─────────┘
↓
┌──────────────────┐
│ Retrieval │
│ + Similarity │
└────────┬─────────┘
↓
┌──────────────────┐
│ Context │
└────────┬─────────┘
↓
┌──────────────────┐
│ LLM │
└────────┬─────────┘
↓
┌──────────────────┐
│ Answer + Sources │
└──────────────────┘
rag-knowledge-assistant/
├── app/
├── data/
│ ├── processed/
│ └── raw/
├── notebooks/
├── src/
│ ├── ingestion.py
│ ├── chunking.py
│ ├── structure_chunking.py
│ ├── semantic_chunking.py
│ ├── chunking_pipeline.py
│ ├── embedding.py
│ ├── retrieval.py
│ └── evaluate_retrieval.py
├── tests/
├── .env.example
├── .gitignore
├── README.md
└── requirements.txt
- Python
- Sentence Transformers
- NumPy
- PyPDF
- Git / GitHub
- Linux
Planned additions include Qdrant, reranking models, FastAPI, Docker, pytest, and GitHub Actions.
Build the core pipeline without relying on RAG frameworks.
Experiment with:
- chunk size and overlap
- semantic search
- metadata filtering
- similarity thresholds
- hybrid retrieval
- reranking
Measure retrieval quality using:
- Recall@K
- Precision@K
- MRR
Then evaluate generated answers for relevance, faithfulness, and citation correctness.
Add:
- FastAPI
- structured logging
- configuration management
- testing
- Docker
- GitHub Actions
- latency measurement
- retrieval and generation tracing
Explore:
- query rewriting
- multi-query retrieval
- HyDE
- parent-child retrieval
- contextual compression
- conversational memory
Advanced techniques will be added only after the baseline system is understood and evaluated.