A local Retrieval-Augmented Generation (RAG) system for PDF documents using LangChain, ChromaDB, and OpenAI-compatible APIs with LM Studio.
- PDF Processing: Extracts text from PDFs and filters out pages with less than 20 characters
- Document Chunking: Splits documents into manageable chunks for better retrieval
- Vector Storage: Uses ChromaDB for efficient similarity search
- Local LLM: Integrates with LM Studio for local language model inference
- Web Interface: Streamlit-based UI for easy interaction
- Batch Processing: Support for processing multiple PDFs or entire directories
-
Activate your virtual environment:
python -m venv .openai_rag_venv
source .openai_rag_venv/bin/activate -
Install Dependencies:
pip install -r requirements.txt
-
Configure LM Studio:
- Start LM Studio and load your
microsoft/phi-4-mini-reasoningmodel - Ensure it's running on
http://localhost:1234 - Update
.envfile if using different settings
- Start LM Studio and load your
-
Run the Application:
streamlit run streamlit_app.py
- Open the web interface
- Test system connectivity
- Upload PDF files or specify a directory
- Ask questions about your documents
from rag_system import RAGSystem
# Initialize the system
rag = RAGSystem()
# Add a PDF
rag.add_pdf("path/to/document.pdf")
# Query the system
result = rag.query("What is the main topic of the document?")
print(result["answer"])Edit .env file to customize:
OPENAI_API_BASE: LM Studio endpoint (default: http://localhost:1234/v1)OPENAI_API_KEY: API key (default: lm-studio)MODEL_NAME: Model name (default: microsoft/phi-4-mini-reasoning)
- PDFProcessor: Handles PDF text extraction and chunking
- VectorStore: Manages ChromaDB operations and embeddings
- LLMClient: Interfaces with LM Studio for text generation
- RAGSystem: Orchestrates the complete RAG pipeline
- Streamlit App: Provides web-based user interface
- Python 3.8+
- LM Studio running locally
- PDF files for processing