A grounded question-answering agent for the portfolio at tts.bedvibe.studio/portfolio.
The agent answers questions using only Panos's published papers and project documentation — it cannot make things up. Every answer includes a "Retrieval X-ray" showing exactly what evidence the retriever fed to the LLM.
Retrieval is documented in
docs/RETRIEVAL.md— backends, per-page scoping, cross-page sessions, and the measured numbers. Decisions are indocs/adr/.Tracing is documented in
obs/README.md— OpenTelemetry spans through the whole answer path plus Langfuse for tokens, cost and sessions. Off unlessOBS_ENABLEDis set; the deployed pipeline is unchanged by it.
Retrieval is easy to build and hard to trust. Most of the work in this repository is about the second part.
The deploy refuses to ship. It will not deploy unless coverage is green and
an authoritative answer proof exists for the exact corpus, questions, prompt and
pinned model — a fingerprint over all four. The proof at the 2026-08-24 deploy was
32/32 questions, coverage 20/20 golden and 112/112 documents, retrieval hit@6 20/20.
See docs/adr/006
and evals/.
A deliberately broken index, to find out whether the gate was real. A
misconfigured approximate index returns plausible results, throws no error, and
finds 14% of the correct neighbours. Answers still looked fine — the keyword
half of the hybrid retriever was quietly carrying it, and they stayed correct
97% of the time. Only recall@k against an exact baseline exposed it.
docs/adr/002
The same benchmark then found two faults nobody planted. Candidate truncation
discarding 13.5% of correct results even with a perfect index — a defect in the
pipeline around the index, not the index. And the Qdrant benchmark was secretly
timing brute force inside the container: a second undocumented threshold made the
query planner ignore the graph it had just built. The measurement tool was
misreporting, and the tool itself is what caught it.
docs/adr/004
Three backends measured, the boring one shipped. Brute force 0.86 ms, FAISS
0.035 ms, Qdrant gRPC 1.3 ms, Qdrant REST 12.4 ms. Transport alone cost 9× the
database. Qdrant held 433 MB to serve a corpus under one megabyte. The
crossover is written down rather than guessed: at roughly one million vectors
FAISS is 302× faster than brute force, and that is where the decision flips.
docs/adr/001
The instrument was wrong before it was right. Its first report was that 34.6%
of retrieval time was spent in a span that does nothing. It was not — the wall
clock ticks in 0.5 ms steps and the work was faster than the tick, so durations
quantised into the wrong span. The clock had to be fixed before the tool could be
trusted. Once honest, it found the real defect: a fresh model client constructed
on every answer, costing more than all retrieval combined. Request p50
12.41 → 2.91 ms, p95 19.35 → 4.45 ms.
docs/adr/008
A 3.5% corpus change broke a tuning constant. Going from 919 to 951 chunks
stopped nprobe=8 reproducing the exact ranking; it had to be raised to 16. Only
the benchmark caught that.
corpus (markdown + HTML)
↓ ingest.py
index/corpus.sqlite (chunks + embeddings)
↓ retrieve.py / scoped_retrieve.py
(hybrid: 0.65 x cosine + 0.35 x BM25, MIN_SIM floor, per-document cap)
↓ llm.py (Gemini Flash)
server.py (FastAPI — POST /api/chat, GET /api/health, GET /api/scopes)
↑
widget/chat.js + chat.css (vanilla JS floating widget)
- Embeddings:
minishlab/potion-base-8M(model2vec static, 256 dims), local CPU, zero API calls - Retrieval: hybrid cosine + BM25 — not cosine-only
- Index: SQLite by default; optional FAISS IVF or Qdrant HNSW behind
RETRIEVAL_BACKEND(seedocs/RETRIEVAL.md) - LLM: Google Gemini Flash via
google-genai - Frontend: vanilla JS + CSS, no frameworks
Corrected 2026-08-22. This section previously said
all-MiniLM-L6-v2embeddings and cosine-only retrieval. Both were wrong against the code (embedder.py:18,retrieve.py:28-31) and the cosine-only error was on record as a known contradiction inorchestrator/PROTOCOL.md. Verified against the source before editing, not assumed.
python -m venv .venv
# Windows
.venv\Scripts\activate
# macOS/Linux
source .venv/bin/activatepip install -r requirements.txtcp .env.example .env
# Edit .env and add your GEMINI_API_KEYpython ingest.pyThe corpus and the site are not in this repository — they are the content
being indexed, not part of the indexer. Point ingest.py at your own sources:
export PORTFOLIO_CORPUS_DIR=/path/to/corpus # markdown papers + projects
export PORTFOLIO_WEB_ROOT=/path/to/site # optional: HTML pages
export PORTFOLIO_BLOG_DIST=/path/to/blog/dist # optional: prerendered blog
python ingest.pyIt chunks by heading, computes embeddings and writes index/corpus.sqlite.
That file is a build artifact and is gitignored. Fourteen tests need it; without
it they skip with a stated reason rather than fail or, worse, pass.
uvicorn server:app --host 0.0.0.0 --port 8000Health check: GET http://localhost:8000/api/health
pytest{
"message": "What is the LongBook Verifier?",
"history": [
{"role": "user", "content": "previous question"},
{"role": "assistant", "content": "previous answer"}
]
}Response:
{
"answer": "The LongBook Verifier is...",
"evidence": [
{
"title": "papers/longbook-verifier/methods.md",
"url": "https://doi.org/...",
"section_path": "Methods > Experiment A",
"chunk_text": "[Methods > Experiment A] ...",
"score": 0.63
}
]
}history: max 6 turnsmessage: max 500 characters- Rate limit: 20 requests/hour per IP
- Empty retrieval → fixed refusal response (no LLM call)
Returns {"status": "ok", "chunks_loaded": "N"}.
Include in any page:
<link rel="stylesheet" href="/widget/chat.css">
<script src="/widget/chat.js"></script>The widget calls POST https://tts.bedvibe.studio/api/chat.
| File | Purpose |
|---|---|
ingest.py |
Corpus chunking + embedding → SQLite index |
retrieve.py |
In-memory hybrid retrieval (cosine + BM25) |
scoped_retrieve.py |
Same ranker, pluggable backend, per-page scope filter |
vectorstore/ |
brute-force / FAISS IVF / Qdrant HNSW backends + scope tags |
session_store.py |
Conversation state that survives a page change |
bench/ |
Recall, latency, scale and span-trace measurement — zero API calls |
obs/ |
OpenTelemetry + Langfuse tracing, cost metering, quota attribution |
docs/RETRIEVAL.md, docs/adr/ |
Retrieval documentation and decisions |
llm.py |
Isolated Gemini API call (swappable) |
server.py |
FastAPI server |
url_map.json |
Human-editable corpus → public URL mapping |
widget/chat.js |
Frontend chat widget |
widget/chat.css |
Widget styles |
tests/ |
pytest suite |