A C++ retrieval-augmented document question-answering system built around custom indexing, multi-term retrieval, and LLM-based grounded synthesis.
DocuQuest indexes a directory of text documents and answers natural-language questions against that collection. An LLM first expands the question into lexical search-term groups; a custom inverted index retrieves documents that satisfy each group; and a second LLM call synthesizes an answer from the selected document text. Retrieval and generation are deliberately separated so the deterministic search layer can be tested without an API key.
flowchart TD
Q[User question] --> E[LLM query expansion]
E --> T[Tokenizer]
T --> I[Inverted index]
I --> R[AND within each term group]
R --> U[Union across term groups]
U --> C[Context selection: up to 10 documents]
C --> S[LLM grounded synthesis]
S --> A[Answer]
Multimap is a custom, unbalanced binary-search tree that maps each normalized token to a vector of unique document paths. It retains the algorithmic behavior of the original implementation while using ownership-aware nodes for deterministic cleanup.
With K distinct keys and V values for a matching key:
| Operation | Average case | Worst case |
|---|---|---|
| Find a key | O(log K) |
O(K) |
| Insert a new key/value pair | O(log K + V) |
O(K + V) |
| Iterate matching values | O(V) |
O(V) |
The linear worst case is important: this tree is not self-balancing.
Tokenizer scans text lazily, emits contiguous ASCII alphanumeric sequences, lowercases alphabetic characters, and treats punctuation or whitespace as delimiters. Indexing and query expansion therefore use the same lexical normalization.
Index associates every token with the documents in which it appears. A multi-term query retrieves the posting list for each term and intersects the lists, so every returned document contains every requested term. Directory traversal and returned paths are sorted for reproducible behavior.
Agent coordinates two model calls around the deterministic index:
- Insert the question into the term-expansion prompt.
- Tokenize each returned line as one AND-query group.
- Union documents found across the groups.
- Read at most ten matching documents in deterministic path order.
- Insert the question and selected text into the synthesis prompt.
- Ask the model for an answer grounded in that context.
The model client is injected through LlmClient, allowing tests to replace the external service with deterministic responses.
.
├── app/
│ └── main.cpp
├── examples/
│ └── sample_docs/
├── include/docuquest/
│ ├── agent.h
│ ├── index.h
│ ├── llm_client.h
│ ├── multimap.h
│ └── tokenizer.h
├── prompts/
│ ├── summarize.txt
│ └── terms.txt
├── src/
│ ├── agent.cpp
│ ├── index.cpp
│ ├── multimap.cpp
│ ├── openrouter_client.cpp
│ └── tokenizer.cpp
├── tests/
│ └── test_core.cpp
├── CMakeLists.txt
└── LICENSE
- CMake 3.20 or newer
- A C++17 compiler
- Windows: WinHTTP (included with the Windows SDK)
- Linux/macOS: libcurl development files
- An OpenRouter API key for live generation only
Deterministic tests do not require network access or credentials.
git clone https://github.com/Stanley-Chow/docuquest-agent.git
cd docuquest-agent
cmake -S . -B build
cmake --build build --config Release
ctest --test-dir build -C Release --output-on-failureCredentials are read only from the process environment. Never place them in the repository.
PowerShell:
$env:OPENROUTER_API_KEY = "your-key"
$env:OPENROUTER_MODEL = "openrouter/auto" # optionalBash/Zsh:
export OPENROUTER_API_KEY="your-key"
export OPENROUTER_MODEL="openrouter/auto" # optionalRun from the repository root so the default prompt paths resolve correctly:
./build/docuquest examples/sample_docsWith a multi-configuration generator on Windows:
.\build\Release\docuquest.exe examples\sample_docsThe first argument may point to any directory of readable text files. Optional second and third arguments override the term-expansion and synthesis prompt files.
The included synthetic corpus supports a small end-to-end demonstration:
Question: Which facility stores the amber navigation charts?
Retrieved context: examples/sample_docs/harbor_archive.txt
Representative grounded answer: The Harbor Archive stores the amber navigation charts.
The retrieved path is determined by the lexical index. Generated wording can vary with the selected model.
- Inverted index over repeated scans: document text is processed once, then term lookups reuse posting lists.
- AND within, OR across expansions: precise multi-term groups reduce false matches, while multiple LLM-generated groups broaden recall.
- Custom multimap: the index exposes the implementation and complexity trade-offs of a purpose-built BST rather than hiding them behind
std::map. - Injected model client: retrieval tests remain deterministic and credential-free.
- Environment-only secrets: no key file or local path is part of the runtime contract.
The portfolio implementation centers on the custom Multimap, Tokenizer, inverted Index, and Agent orchestration. The current public release uses independent application, interface, build, and OpenRouter-adapter code; externally supplied framework utilities, driver code, diagrams, and document archives are not part of the release tree.
- Retrieval is lexical; there are no embeddings, semantic reranker, stemming, or phrase positions.
- The BST is not balanced and can degrade to linear lookup time for adversarial insertion order.
- Context selection is deterministic path order rather than learned relevance ranking.
- Documents are included whole, up to a configurable cap, rather than chunked to a token budget.
- Query expansion and answer quality depend on the selected external model.
- Grounding prompts reduce unsupported output but cannot guarantee that a model will not hallucinate.
Released under the MIT License.