A Retrieval-Augmented Generation (RAG) system built in pure Python — no
LangChain, no LlamaIndex. Dependencies are managed with uv,
document retrieval currently runs on ElasticSearch, and generation goes through
Anthropic, OpenAI or Groq — selectable at runtime.
The pipeline is three independent, testable stages, orchestrated by a single
rag(query) function:
query --> search(query) --> build_prompt(query, results) --> llm(prompt) --> answer
src/search.py src/prompt.py src/llm.py
- Search (
src/search.py) — retrieves the top-N relevant documents from the index for a given query. - Prompt (
src/prompt.py) — formats the retrieved context and the question into the LLM's prompt template. - LLM (
src/llm.py) — sends the prompt to the selected provider and returns the generated answer.
Roadmap: the ElasticSearch backend is planned to be replaced by
PostgreSQL + pgvector in production, to simplify disaster recovery
(physical/logical replication, dumps) and to get ACID guarantees on
ingestion.
- Python >= 3.12
uv- Docker (for local ElasticSearch)
# Install dependencies and create the virtual environment
uv sync
# Configure environment variables
cp .env.example .env
# then fill in the key for the provider you intend to use
# Start ElasticSearch locally
docker-compose up -d elasticsearchIdentity-linked API keys must also set
ANTHROPIC_WORKSPACE_ID. The API rejects them with a400 invalid_request_error— "anthropic-workspace-id is required" — unless the request names the workspace it acts in. The id is in the Anthropic Console under Settings → Workspaces, and appears in the URL aswrkspc_.... A standard Console API key needs no such header, so leave the variable blank for one.
If
uvor Docker fail TLS/certificate validation behind a corporate proxy or antivirus that intercepts HTTPS, seeuv.toml(system-certs = true) — it's already configured to use the OS certificate store instead ofuv's bundled one.
src/cli.py is the entrypoint that drives the pipeline from the terminal. It
is registered as the rag console script by uv sync, and also runs as
python -m src.cli.
# 1. Load a corpus (a JSON file: a list of {"title", "content"} objects)
uv run rag ingest data/sample_documents.json
# 2. Inspect retrieval on its own — no LLM call, so no API cost
uv run rag search "When are deploys frozen?"
# 3. Run the full pipeline and get a generated answer
uv run rag ask "When are deploys frozen?"search is deliberately separate from ask: when an answer comes back wrong,
retrieval is usually what is wrong, and diagnosing that should not cost an API
call. --size controls how many documents come back, and --index overrides
the target index for a single run.
The index is read from ELASTICSEARCH_INDEX (default rag_docs), so ingest
and ask always agree on where the corpus lives:
ELASTICSEARCH_INDEX=my_corpus uv run rag ingest my_documents.json
ELASTICSEARCH_INDEX=my_corpus uv run rag ask "What does it say?"Exit codes are 0 on success, 1 on a runtime failure (unreachable
ElasticSearch, a missing provider key, malformed corpus file) and 2 on a
usage error, so the commands compose in a shell script.
Three providers are available: anthropic (default), openai and groq.
LLM_PROVIDER sets the deployment's default and --provider overrides it for
one question, so the same question can be put to each of them without editing
anything:
uv run rag ask --provider openai "When are deploys frozen?"
uv run rag ask --provider groq "When are deploys frozen?"Groq is not a third integration. It speaks the OpenAI wire format, so it reuses that SDK against a different base URL rather than adding a dependency for the same request shape.
Only the selected provider's key is required — ANTHROPIC_API_KEY,
OPENAI_API_KEY or GROQ_API_KEY, each the name that provider's own SDK
already expects. A missing one is reported before retrieval runs, naming the
variable the chosen provider actually needs.
The SDKs disagree on how they report a refusal and a truncation — Anthropic
uses stop_reason, OpenAI uses finish_reason plus a separate refusal
field — and that difference stops inside src/llm.py. Every provider reaches
the pipeline as one of the same three outcomes: an answer, LLMRefusalError,
or LLMTruncatedError.
Each provider carries its own default and its own override variable, so switching provider does not silently carry the previous provider's model name along:
| Provider | Variable | Default |
|---|---|---|
anthropic |
ANTHROPIC_MODEL |
claude-sonnet-5 |
openai |
OPENAI_MODEL |
gpt-4o |
groq |
GROQ_MODEL |
llama-3.3-70b-versatile |
ANTHROPIC_MODEL=claude-opus-5 uv run rag ask "Something genuinely hard"These defaults are a starting point, not a guarantee that your account has access to that particular model — set the variable if it does not.
LLM_MAX_TOKENS (default 4096) bounds the generated answer on every
provider. It is a ceiling, not a budget: only tokens actually generated are
billed, so raising it costs nothing by itself. It exists to stop a runaway
generation.
Thinking is enabled by default on current models and its tokens count against the same ceiling, so a value sized for the answer alone would truncate routinely. An answer that does hit the ceiling raises rather than returning what it got — a half-finished answer is indistinguishable from a complete one by looking at the text, which in a RAG system is worse than an error.
This project follows BDD/TDD: every feature starts as a Gherkin spec in
tests/features/, bound to step definitions in tests/step_defs/ via
pytest-bdd, with pytest covering isolated logic in tests/unit/.
# Full suite
uv run pytest
# Verbose output
uv run pytest -vEverything under tests/step_defs/ and tests/unit/ mocks the ElasticSearch
client, so none of it would catch a client/server incompatibility.
tests/integration/ talks to a real server and is the only place where the
pinned client is actually verified against the running cluster.
These tests skip automatically when no server is reachable, so the plain
uv run pytest above works without Docker. To run them for real:
docker-compose up -d elasticsearch
uv run pytest -m integration -vClient and server must stay on the same major version — the pin in
pyproject.toml (elasticsearch>=8.15.0,<9.0.0) and the image tag in
docker-compose.yml are two halves of one decision, and
test_client_and_server_majors_are_compatible fails loudly if they drift.
# Build the app image and run the full suite against a live ElasticSearch
docker-compose up --build
# Just the database, for local development
docker-compose up -d elasticsearchThe app service waits for the elasticsearch healthcheck before starting,
and its default command runs the test suite.
.
├── tests/
│ ├── features/ # Gherkin (.feature) specs
│ ├── step_defs/ # pytest-bdd step definitions
│ ├── unit/ # Unit tests
│ └── integration/ # Tests against a real ElasticSearch
├── src/
│ ├── search.py # ElasticSearch integration (future: PostgreSQL)
│ ├── prompt.py # Context/question prompt formatting
│ ├── llm.py # LLM API calls (Anthropic / OpenAI / Groq)
│ ├── rag.py # Orchestrates search -> prompt -> llm
│ └── cli.py # Terminal entrypoint (ingest / search / ask)
├── data/
│ └── sample_documents.json # Small corpus for trying the CLI out
├── Dockerfile # Application image
├── docker-compose.yml # ElasticSearch + app
├── .env.example # Environment variable template
└── pyproject.toml # uv/pytest configuration
Progress is tracked as GitHub Issues against a BDD backlog, organized on the
RAG PoC Roadmap project board.
Every story starts as a Gherkin spec, so the acceptance criteria on an issue
and the .feature file that verifies it are the same text.
| User story | Spec | Status | |
|---|---|---|---|
| US01 | Automatically create the ElasticSearch index on startup | index_initialization.feature |
Done |
| US02 | Bulk-ingest documents into the knowledge base | ingestion.feature |
Done |
| US03 | Search for relevant documents by query | search.feature |
Done |
| US04 | Format prompt with retrieved context and question | prompt_formatting.feature |
Done |
| US05 | Generate LLM answer from formatted prompt | answer_generation.feature |
Done |
| US06 | Orchestrate the full RAG flow in rag(query) |
rag_orchestration.feature |
Done |
| US07 | Drive the pipeline end to end from the command line | cli.feature |
Done |
| US08 | Answer through Anthropic, OpenAI or Groq | provider_selection.feature |
Done |
One link in the chain has never run against a live API: a successful response
from any provider. Retrieval, prompt construction, orchestration and provider
selection have all been exercised against a real ElasticSearch, but the
generation step is covered by stubs shaped from the documented response
formats. Supplying a working key and running rag ask closes it.
- Ingest a real corpus rather than the six-document sample
- Replace ElasticSearch with PostgreSQL +
pgvector(see Roadmap above)
Distributed under the MIT License.