Skip to content

Latest commit

 

History

13 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Enterprise Multilingual RAG Platform

Python Version FastAPI Groq LPU Embedding Model Vector Store License: MIT

A production-grade, enterprise Multilingual Retrieval-Augmented Generation (RAG) platform. Ingest heterogeneous documents (PDF, DOCX, PPTX, XLSX, TXT), perform cross-lingual semantic vector search across 100+ languages, and generate grounded answers with precise source citations.

Features • System Architecture • Quick Start • API Specification • Supported Formats • Testing


Overview

The Enterprise Multilingual RAG Platform provides a scalable, modular REST API for enterprise question-answering over diverse document collections. It solves the challenge of cross-lingual knowledge retrieval by combining BAAI/bge-m3 (state-of-the-art multilingual embedding model covering 100+ languages) with ChromaDB local vector storage and high-throughput inference powered by Groq (llama-3.3-70b-versatile).

Every synthesized answer includes deterministic source citations specifying the exact source document, page number, chunk index, and relevance score, strictly preventing hallucinations.


Key Features

  • 🌐 True Multilingual Retrieval: Powered by BAAI/bge-m3, enabling semantic similarity search across more than 100 languages. Supports cross-lingual queries (e.g., asking an English question over a Spanish, German, French, Arabic, or Hindi document).
  • 📄 Multi-Format Document Ingestion: Built-in parsers for PDF (PyMuPDF), Word (docx), PowerPoint (pptx), Excel (xlsx), and plain Text (txt).
  • 🧩 Context-Aware Semantic Chunking: Recursive token-aware character chunking with configurable chunk sizes and overlaps, preserving critical document metadata (document ID, page numbers, chunk index, language).
  • ⚡ Ultra-Fast Inference via Groq LPU: Leverages Groq's high-speed inference engine running llama-3.3-70b-versatile to deliver sub-second generation latencies.
  • 🎯 Grounded Answers with Source Citations: Synthesized answers return human-readable citations (e.g., [Source: quarterly_report.pdf, Page 4]) linked directly to the retrieved vector chunks.
  • 🛡️ Anti-Hallucination Guardrails: Specialized system prompt instructs the model to rely strictly on retrieved context. If information is absent from the ingested corpus, the system explicitly states it cannot answer.
  • 🚀 Production-Ready REST API: Clean FastAPI architecture with Swagger UI documentation, strict Pydantic v2 schemas, automated health checks, and dependency injection.

System Architecture

                       ┌─────────────────────────┐
                       │  User / Client App      │
                       └────────────┬────────────┘
                                    │
           ┌────────────────────────┴────────────────────────┐
           │                                                 │
 [ Ingestion Flow: POST /upload + /ingest ]       [ Query Flow: POST /ask ]
           │                                                 │
           ▼                                                 ▼
┌───────────────────────┐                         ┌───────────────────────┐
│ Multi-Format Parsers  │                         │ User Question         │
│ (PDF, DOCX, PPTX,     │                         │ (Any Language)        │
│  XLSX, TXT)           │                         └───────────┬───────────┘
└──────────┬────────────┘                                     │
           ▼                                                  ▼
┌───────────────────────┐                         ┌───────────────────────┐
│ Chunking Service      │                         │ Embedding Service     │
│ (Recursive Splitter,  │                         │ (BAAI/bge-m3 dense    │
│  Metadata Tracking)   │                         │  vector generation)   │
└──────────┬────────────┘                         └───────────┬───────────┘
           ▼                                                  │
┌───────────────────────┐                                     │
│ Embedding Service     │                                     ▼
│ (BAAI/bge-m3 1024-dim)│                         ┌───────────────────────┐
└──────────┬────────────┘                         │ ChromaDB Vector Search│
           ▼                                      │ (Cosine Similarity &  │
┌───────────────────────┐                         │  Metadata Filtering)  │
│ ChromaDB Vector Store │◄────────────────────────┴───────────┬───────────┘
│ (Persistent Storage)  │                                     │
└───────────────────────┘                         [ Top-K Ranked Chunks ]
                                                              │
                                                              ▼
                                                  ┌───────────────────────┐
                                                  │ Context & Prompt      │
                                                  │ Assembly              │
                                                  └───────────┬───────────┘
                                                              ▼
                                                  ┌───────────────────────┐
                                                  │ Groq LLM Generation   │
                                                  │ (Llama 3.3 70B)       │
                                                  └───────────┬───────────┘
                                                              ▼
                                                  ┌───────────────────────┐
                                                  │ Answer + Citations    │
                                                  │ (Page, Source, Score) │
                                                  └───────────────────────┘

Supported Document Formats

Format Extension Engine / Library Extracted Metadata
Adobe PDF .pdf PyMuPDF (fitz) Filename, Page Number, Chunk Index
Microsoft Word .docx python-docx Filename, Paragraph / Section, Chunk Index
Microsoft PowerPoint .pptx python-pptx Filename, Slide Number, Chunk Index
Microsoft Excel .xlsx openpyxl Filename, Sheet Name, Row/Cell Data
Plain Text / Markdown .txt, .md Built-in UTF-8 parser Filename, Chunk Index

Tech Stack

Component Technology Description
API Framework FastAPI Async REST API framework with Swagger docs
Runtime & Server Uvicorn Lightning-fast ASGI web server
Validation Pydantic v2 Strict data parsing and schema validation
LLM Inference Groq LPU inference running llama-3.3-70b-versatile
Embedding Model BAAI/bge-m3 1024-dimension dense multilingual embeddings
Vector Database ChromaDB Local embedded vector database with persistent storage
Document Processing PyMuPDF, python-docx, python-pptx, openpyxl Robust document text extraction
Chunking LangChain Recursive character text splitting

Project Structure

├── app/
│   ├── main.py                       # FastAPI application factory & router inclusion
│   ├── config.py                     # Centralized Pydantic application settings
│   ├── api/
│   │   ├── dependencies.py           # Dependency injection providers for services
│   │   └── routes/
│   │       ├── health.py             # Health check endpoint (/health)
│   │       ├── upload.py             # Document upload endpoint (/upload)
│   │       ├── ingest.py             # Ingestion & vectorization endpoint (/ingest)
│   │       └── query.py              # Question-answering endpoint (/ask)
│   ├── parsers/
│   │   ├── base.py                   # Base parser interface
│   │   ├── pdf_parser.py             # PDF extraction with page tracking
│   │   ├── docx_parser.py            # Word document parser
│   │   ├── pptx_parser.py            # PowerPoint slide parser
│   │   ├── xlsx_parser.py            # Excel spreadsheet parser
│   │   ├── txt_parser.py             # Plain text parser
│   │   └── registry.py               # Dynamic parser factory
│   ├── chunking/
│   │   └── chunking_service.py       # Recursive text splitting & metadata mapping
│   ├── embeddings/
│   │   └── embedding_service.py      # BAAI/bge-m3 embedding generator
│   ├── vectorstore/
│   │   └── vectorstore_service.py    # ChromaDB collection management & search
│   ├── retrieval/
│   │   └── retrieval_service.py      # Semantic similarity retrieval & filtering
│   ├── prompts/
│   │   └── rag_prompt.py             # Anti-hallucination prompt templates
│   ├── schemas/                      # Pydantic request & response models
│   │   ├── answer.py                 # GeneratedAnswer & Citation models
│   │   ├── chunk.py                  # Text chunk data models
│   │   ├── embedding.py              # Embedding vector models
│   │   ├── parsed_document.py        # Parsed document schemas
│   │   ├── upload.py                 # Upload metadata models
│   │   └── vectorstore.py            # Search result models
│   └── services/
│       ├── answer_generation.py      # Groq LLM answer synthesizer
│       ├── file_storage.py           # Disk storage management for uploads
│       └── rag_pipeline.py           # Unified orchestrator (Query -> Retrieval -> LLM)
├── documents/                        # Local document storage (.gitkeep)
├── tests/                            # Comprehensive automated test suite
│   ├── test_parsers.py               # Parser unit tests
│   ├── test_chunking.py              # Chunking unit tests
│   ├── test_embeddings.py            # Embedding service tests
│   ├── test_vectorstore.py           # ChromaDB vector store tests
│   ├── test_retrieval.py             # Retrieval service tests
│   ├── test_answer_generation.py     # Answer generation & citation tests
│   ├── test_rag_pipeline.py          # RAG pipeline tests
│   └── test_e2e_pipeline.py          # Full end-to-end integration test
├── .env.example                      # Configuration template
├── .gitignore                        # Git exclusion rules
├── main.py                           # Application startup script
├── pyproject.toml / pyrefly.toml     # Linter & project configuration
└── requirements.txt                  # Python dependencies

Quick Start

1. Prerequisites

  • Python: 3.10 or higher
  • Groq API Key: Free API key from Groq Console

2. Clone the Repository & Set Up Virtual Environment

git clone https://github.com/Afaque-Khan04/Multilingual-RAG.git
cd Multilingual-RAG

# Create virtual environment
python -m venv .venv

# Activate virtual environment
# Windows (PowerShell):
.\.venv\Scripts\Activate.ps1
# Linux / macOS:
source .venv/bin/activate

# Install dependencies
pip install -r requirements.txt

3. Configure Environment Variables

Copy .env.example to .env and set your Groq API key:

cp .env.example .env

Edit .env:

APP_NAME=Multilingual RAG API
APP_ENV=development
DEBUG=true
HOST=0.0.0.0
PORT=8000

# Groq API Key
GROQ_API_KEY=gsk_your_groq_api_key_here

# Suppress Hugging Face symlink warning on Windows
HF_HUB_DISABLE_SYMLINKS_WARNING=1

4. Run the API Server

# Using uvicorn directly
uvicorn app.main:app --reload --port 8000

# Or via main.py
python main.py

The server will start at http://127.0.0.1:8000.

Open Interactive Swagger UI:

http://127.0.0.1:8000/docs

API Specification

Endpoints Overview

Method Endpoint Description Request Response
GET /health System health check None {"status": "ok"}
POST /upload Upload a document to storage multipart/form-data UploadMetadata
POST /ingest Parse, chunk, embed & index document JSON (document_id, language) IngestResponse
POST /ask Query the RAG pipeline JSON (question, top_k) GeneratedAnswer

Step-by-Step Usage Examples

1. Upload a Document

curl -X POST "http://127.0.0.1:8000/upload" \
  -F "file=@sample_report.pdf"

Response:

{
  "id": "a4bf5ca3-2911-47e3-a517-169d9e7b9e07",
  "original_filename": "sample_report.pdf",
  "content_type": "application/pdf",
  "size_bytes": 177656,
  "storage_path": "documents/a4bf5ca3-2911-47e3-a517-169d9e7b9e07.pdf",
  "uploaded_at": "2026-09-08T10:30:00Z"
}

2. Ingest the Document into ChromaDB

curl -X POST "http://127.0.0.1:8000/ingest" \
  -H "Content-Type: application/json" \
  -d '{
    "document_id": "a4bf5ca3-2911-47e3-a517-169d9e7b9e07",
    "language": "en"
  }'

Response:

{
  "document_id": "a4bf5ca3-2911-47e3-a517-169d9e7b9e07",
  "filename": "sample_report.pdf",
  "chunks_created": 14,
  "message": "Successfully ingested 14 chunks into ChromaDB."
}

3. Ask a Question

curl -X POST "http://127.0.0.1:8000/ask" \
  -H "Content-Type: application/json" \
  -d '{
    "question": "What were the primary revenue drivers in Q3?",
    "top_k": 3
  }'

Response:

{
  "answer": "The primary revenue drivers in Q3 were enterprise cloud subscriptions and international expansion in European markets, representing a 28% increase year-over-year.",
  "query": "What were the primary revenue drivers in Q3?",
  "citations": [
    {
      "source_filename": "sample_report.pdf",
      "page_number": 3,
      "document_id": "a4bf5ca3-2911-47e3-a517-169d9e7b9e07",
      "chunk_index": 2,
      "score": 0.892
    }
  ],
  "formatted_citations": [
    "[Source: sample_report.pdf, Page 3]"
  ],
  "model": "llama-3.3-70b-versatile",
  "context_used": 3,
  "answer_with_citations": "The primary revenue drivers in Q3 were enterprise cloud subscriptions and international expansion in European markets, representing a 28% increase year-over-year.\n\nSources:\n- [Source: sample_report.pdf, Page 3]"
}

Configuration Reference

All settings can be configured via environment variables or the .env file:

Setting Default Description
APP_NAME Multilingual RAG API Application title displayed in OpenAPI docs
DEBUG true Enable debug mode and automatic reload
HOST 0.0.0.0 Bind host address
PORT 8000 Port number for web server
GROQ_API_KEY "" API Key for Groq LPU inference
LLM_MODEL llama-3.3-70b-versatile LLM model name on Groq
LLM_TEMPERATURE 0.1 Sampling temperature (low value reduces hallucinations)
LLM_MAX_TOKENS 1024 Maximum tokens in generated response
EMBEDDING_MODEL BAAI/bge-m3 Hugging Face model for multilingual embeddings
EMBEDDING_BATCH_SIZE 32 Batch size during vectorization
CHUNK_SIZE 1000 Maximum character length per chunk
CHUNK_OVERLAP 200 Character overlap between contiguous chunks
CHROMADB_PERSIST_DIR ./chromadb_data Directory for persistent ChromaDB storage
RETRIEVAL_TOP_K 5 Default number of vector chunks retrieved per query

Running Tests

The test suite covers unit tests for parsers, chunking, embeddings, vector store, and full end-to-end pipeline execution:

# Run all unit tests
pytest -v

# Run specific component tests
pytest tests/test_parsers.py -v
pytest tests/test_chunking.py -v
pytest tests/test_embeddings.py -v
pytest tests/test_vectorstore.py -v
pytest tests/test_answer_generation.py -v

# Run End-to-End integration test against a running server
python tests/test_e2e_pipeline.py

Contributing

Contributions are welcome! To contribute:

  1. Fork the repository.
  2. Create a descriptive feature branch (git checkout -b feature/AmazingFeature).
  3. Commit your changes (git commit -m 'feat: Add AmazingFeature').
  4. Push to the branch (git push origin feature/AmazingFeature).
  5. Open a Pull Request.

About

Production-grade Multilingual RAG API powered by FastAPI, BAAI/bge-m3 embeddings (100+ languages), ChromaDB, and Groq (Llama 3.3 70B). Features multi-format ingestion (PDF, DOCX, PPTX, XLSX, TXT) and grounded answers with deterministic source citations.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages