Skip to content

Repository files navigation

Portfolio RAG Chat Agent

A grounded question-answering agent for the portfolio at tts.bedvibe.studio/portfolio.

The agent answers questions using only Panos's published papers and project documentation — it cannot make things up. Every answer includes a "Retrieval X-ray" showing exactly what evidence the retriever fed to the LLM.

Retrieval is documented in docs/RETRIEVAL.md — backends, per-page scoping, cross-page sessions, and the measured numbers. Decisions are in docs/adr/.

Tracing is documented in obs/README.md — OpenTelemetry spans through the whole answer path plus Langfuse for tokens, cost and sessions. Off unless OBS_ENABLED is set; the deployed pipeline is unchanged by it.

What is worth reading here

Retrieval is easy to build and hard to trust. Most of the work in this repository is about the second part.

The deploy refuses to ship. It will not deploy unless coverage is green and an authoritative answer proof exists for the exact corpus, questions, prompt and pinned model — a fingerprint over all four. The proof at the 2026-08-24 deploy was 32/32 questions, coverage 20/20 golden and 112/112 documents, retrieval hit@6 20/20. See docs/adr/006 and evals/.

A deliberately broken index, to find out whether the gate was real. A misconfigured approximate index returns plausible results, throws no error, and finds 14% of the correct neighbours. Answers still looked fine — the keyword half of the hybrid retriever was quietly carrying it, and they stayed correct 97% of the time. Only recall@k against an exact baseline exposed it. docs/adr/002

The same benchmark then found two faults nobody planted. Candidate truncation discarding 13.5% of correct results even with a perfect index — a defect in the pipeline around the index, not the index. And the Qdrant benchmark was secretly timing brute force inside the container: a second undocumented threshold made the query planner ignore the graph it had just built. The measurement tool was misreporting, and the tool itself is what caught it. docs/adr/004

Three backends measured, the boring one shipped. Brute force 0.86 ms, FAISS 0.035 ms, Qdrant gRPC 1.3 ms, Qdrant REST 12.4 ms. Transport alone cost 9× the database. Qdrant held 433 MB to serve a corpus under one megabyte. The crossover is written down rather than guessed: at roughly one million vectors FAISS is 302× faster than brute force, and that is where the decision flips. docs/adr/001

The instrument was wrong before it was right. Its first report was that 34.6% of retrieval time was spent in a span that does nothing. It was not — the wall clock ticks in 0.5 ms steps and the work was faster than the tick, so durations quantised into the wrong span. The clock had to be fixed before the tool could be trusted. Once honest, it found the real defect: a fresh model client constructed on every answer, costing more than all retrieval combined. Request p50 12.41 → 2.91 ms, p95 19.35 → 4.45 ms. docs/adr/008

A 3.5% corpus change broke a tuning constant. Going from 919 to 951 chunks stopped nprobe=8 reproducing the exact ranking; it had to be raised to 16. Only the benchmark caught that.

Architecture

corpus (markdown + HTML)
    ↓ ingest.py
index/corpus.sqlite (chunks + embeddings)
    ↓ retrieve.py / scoped_retrieve.py
      (hybrid: 0.65 x cosine + 0.35 x BM25, MIN_SIM floor, per-document cap)
    ↓ llm.py (Gemini Flash)
server.py (FastAPI — POST /api/chat, GET /api/health, GET /api/scopes)
    ↑
widget/chat.js + chat.css (vanilla JS floating widget)
  • Embeddings: minishlab/potion-base-8M (model2vec static, 256 dims), local CPU, zero API calls
  • Retrieval: hybrid cosine + BM25 — not cosine-only
  • Index: SQLite by default; optional FAISS IVF or Qdrant HNSW behind RETRIEVAL_BACKEND (see docs/RETRIEVAL.md)
  • LLM: Google Gemini Flash via google-genai
  • Frontend: vanilla JS + CSS, no frameworks

Corrected 2026-08-22. This section previously said all-MiniLM-L6-v2 embeddings and cosine-only retrieval. Both were wrong against the code (embedder.py:18, retrieve.py:28-31) and the cosine-only error was on record as a known contradiction in orchestrator/PROTOCOL.md. Verified against the source before editing, not assumed.

Quick Start

1. Create a virtual environment

python -m venv .venv
# Windows
.venv\Scripts\activate
# macOS/Linux
source .venv/bin/activate

2. Install dependencies

pip install -r requirements.txt

3. Set up environment

cp .env.example .env
# Edit .env and add your GEMINI_API_KEY

4. Build the index

python ingest.py

The corpus and the site are not in this repository — they are the content being indexed, not part of the indexer. Point ingest.py at your own sources:

export PORTFOLIO_CORPUS_DIR=/path/to/corpus        # markdown papers + projects
export PORTFOLIO_WEB_ROOT=/path/to/site            # optional: HTML pages
export PORTFOLIO_BLOG_DIST=/path/to/blog/dist      # optional: prerendered blog
python ingest.py

It chunks by heading, computes embeddings and writes index/corpus.sqlite. That file is a build artifact and is gitignored. Fourteen tests need it; without it they skip with a stated reason rather than fail or, worse, pass.

5. Start the server

uvicorn server:app --host 0.0.0.0 --port 8000

Health check: GET http://localhost:8000/api/health

6. Run tests

pytest

API

POST /api/chat

{
  "message": "What is the LongBook Verifier?",
  "history": [
    {"role": "user", "content": "previous question"},
    {"role": "assistant", "content": "previous answer"}
  ]
}

Response:

{
  "answer": "The LongBook Verifier is...",
  "evidence": [
    {
      "title": "papers/longbook-verifier/methods.md",
      "url": "https://doi.org/...",
      "section_path": "Methods > Experiment A",
      "chunk_text": "[Methods > Experiment A] ...",
      "score": 0.63
    }
  ]
}
  • history: max 6 turns
  • message: max 500 characters
  • Rate limit: 20 requests/hour per IP
  • Empty retrieval → fixed refusal response (no LLM call)

GET /api/health

Returns {"status": "ok", "chunks_loaded": "N"}.

Widget

Include in any page:

<link rel="stylesheet" href="/widget/chat.css">
<script src="/widget/chat.js"></script>

The widget calls POST https://tts.bedvibe.studio/api/chat.

File Map

File Purpose
ingest.py Corpus chunking + embedding → SQLite index
retrieve.py In-memory hybrid retrieval (cosine + BM25)
scoped_retrieve.py Same ranker, pluggable backend, per-page scope filter
vectorstore/ brute-force / FAISS IVF / Qdrant HNSW backends + scope tags
session_store.py Conversation state that survives a page change
bench/ Recall, latency, scale and span-trace measurement — zero API calls
obs/ OpenTelemetry + Langfuse tracing, cost metering, quota attribution
docs/RETRIEVAL.md, docs/adr/ Retrieval documentation and decisions
llm.py Isolated Gemini API call (swappable)
server.py FastAPI server
url_map.json Human-editable corpus → public URL mapping
widget/chat.js Frontend chat widget
widget/chat.css Widget styles
tests/ pytest suite

About

Grounded RAG agent: hybrid retrieval, a deploy gate that refuses to ship without an answer proof, and the planted faults that proved the gate was real

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages