Skip to content

Repository files navigation

DocuQuest Agent

A C++ retrieval-augmented document question-answering system built around custom indexing, multi-term retrieval, and LLM-based grounded synthesis.

Overview

DocuQuest indexes a directory of text documents and answers natural-language questions against that collection. An LLM first expands the question into lexical search-term groups; a custom inverted index retrieves documents that satisfy each group; and a second LLM call synthesizes an answer from the selected document text. Retrieval and generation are deliberately separated so the deterministic search layer can be tested without an API key.

Architecture

flowchart TD
    Q[User question] --> E[LLM query expansion]
    E --> T[Tokenizer]
    T --> I[Inverted index]
    I --> R[AND within each term group]
    R --> U[Union across term groups]
    U --> C[Context selection: up to 10 documents]
    C --> S[LLM grounded synthesis]
    S --> A[Answer]
Loading

Core Components

Multimap

Multimap is a custom, unbalanced binary-search tree that maps each normalized token to a vector of unique document paths. It retains the algorithmic behavior of the original implementation while using ownership-aware nodes for deterministic cleanup.

With K distinct keys and V values for a matching key:

Operation Average case Worst case
Find a key O(log K) O(K)
Insert a new key/value pair O(log K + V) O(K + V)
Iterate matching values O(V) O(V)

The linear worst case is important: this tree is not self-balancing.

Tokenizer

Tokenizer scans text lazily, emits contiguous ASCII alphanumeric sequences, lowercases alphabetic characters, and treats punctuation or whitespace as delimiters. Indexing and query expansion therefore use the same lexical normalization.

Inverted Index

Index associates every token with the documents in which it appears. A multi-term query retrieves the posting list for each term and intersects the lists, so every returned document contains every requested term. Directory traversal and returned paths are sorted for reproducible behavior.

Agent

Agent coordinates two model calls around the deterministic index:

  1. Insert the question into the term-expansion prompt.
  2. Tokenize each returned line as one AND-query group.
  3. Union documents found across the groups.
  4. Read at most ten matching documents in deterministic path order.
  5. Insert the question and selected text into the synthesis prompt.
  6. Ask the model for an answer grounded in that context.

The model client is injected through LlmClient, allowing tests to replace the external service with deterministic responses.

Project Structure

.
├── app/
│   └── main.cpp
├── examples/
│   └── sample_docs/
├── include/docuquest/
│   ├── agent.h
│   ├── index.h
│   ├── llm_client.h
│   ├── multimap.h
│   └── tokenizer.h
├── prompts/
│   ├── summarize.txt
│   └── terms.txt
├── src/
│   ├── agent.cpp
│   ├── index.cpp
│   ├── multimap.cpp
│   ├── openrouter_client.cpp
│   └── tokenizer.cpp
├── tests/
│   └── test_core.cpp
├── CMakeLists.txt
└── LICENSE

Build and Run

Prerequisites

  • CMake 3.20 or newer
  • A C++17 compiler
  • Windows: WinHTTP (included with the Windows SDK)
  • Linux/macOS: libcurl development files
  • An OpenRouter API key for live generation only

Deterministic tests do not require network access or credentials.

Build

git clone https://github.com/Stanley-Chow/docuquest-agent.git
cd docuquest-agent
cmake -S . -B build
cmake --build build --config Release
ctest --test-dir build -C Release --output-on-failure

Configure credentials

Credentials are read only from the process environment. Never place them in the repository.

PowerShell:

$env:OPENROUTER_API_KEY = "your-key"
$env:OPENROUTER_MODEL = "openrouter/auto"  # optional

Bash/Zsh:

export OPENROUTER_API_KEY="your-key"
export OPENROUTER_MODEL="openrouter/auto"  # optional

Start the CLI

Run from the repository root so the default prompt paths resolve correctly:

./build/docuquest examples/sample_docs

With a multi-configuration generator on Windows:

.\build\Release\docuquest.exe examples\sample_docs

The first argument may point to any directory of readable text files. Optional second and third arguments override the term-expansion and synthesis prompt files.

Example

The included synthetic corpus supports a small end-to-end demonstration:

Question: Which facility stores the amber navigation charts?
Retrieved context: examples/sample_docs/harbor_archive.txt
Representative grounded answer: The Harbor Archive stores the amber navigation charts.

The retrieved path is determined by the lexical index. Generated wording can vary with the selected model.

Design Decisions

  • Inverted index over repeated scans: document text is processed once, then term lookups reuse posting lists.
  • AND within, OR across expansions: precise multi-term groups reduce false matches, while multiple LLM-generated groups broaden recall.
  • Custom multimap: the index exposes the implementation and complexity trade-offs of a purpose-built BST rather than hiding them behind std::map.
  • Injected model client: retrieval tests remain deterministic and credential-free.
  • Environment-only secrets: no key file or local path is part of the runtime contract.

Implementation Scope

The portfolio implementation centers on the custom Multimap, Tokenizer, inverted Index, and Agent orchestration. The current public release uses independent application, interface, build, and OpenRouter-adapter code; externally supplied framework utilities, driver code, diagrams, and document archives are not part of the release tree.

Limitations

  • Retrieval is lexical; there are no embeddings, semantic reranker, stemming, or phrase positions.
  • The BST is not balanced and can degrade to linear lookup time for adversarial insertion order.
  • Context selection is deterministic path order rather than learned relevance ranking.
  • Documents are included whole, up to a configurable cap, rather than chunked to a token budget.
  • Query expansion and answer quality depend on the selected external model.
  • Grounding prompts reduce unsupported output but cannot guarantee that a model will not hallucinate.

License

Released under the MIT License.

About

C++ retrieval-augmented document QA with custom indexing, multi-term retrieval, and LLM-based synthesis.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages