From 1a0462add83299c4239d43f38bce839c4a4dd0e9 Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 25 Aug 2026 00:34:40 +0000 Subject: [PATCH] docs: add root CLAUDE.md, overhaul BENCHMARK.md, add CODE_HEALTH.md MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - CLAUDE.md (new): commands, architecture map + layering rules, conventions, testing reality, benchmark quick-ref, branch/PR rules, gotchas. Two docs already declared themselves "extensions of CLAUDE.md" — it now exists. - docs/BENCHMARK.md: remove 5 dead links to nonexistent docs/benchmark/, align File Map/status tables with the 9 real benchmark dirs (7 retired ones removed), fix the --event-stream guidance (flag exists only in locomo; other runners consolidate unconditionally in base_runner), document longmemeval entry points, correct the auto-download claim, add an environment-variable table and a "Contributing a new benchmark" guide. - docs/CODE_HEALTH.md (new): evidence-based structural findings (file:line) in three tiers with fix directions and sequencing. - .env.example: document benchmark env vars; fix duplicate COGNIFOLD_SESSION_BACKEND definition. Docs-only change; no src/ code touched. Co-Authored-By: Claude Claude-Session: https://claude.ai/code/session_01VnAaidzJkganncouErUp3m --- .env.example | 29 +++- CLAUDE.md | 152 ++++++++++++++++ docs/BENCHMARK.md | 412 +++++++++++++++++++------------------------- docs/CODE_HEALTH.md | 253 +++++++++++++++++++++++++++ 4 files changed, 614 insertions(+), 232 deletions(-) create mode 100644 CLAUDE.md create mode 100644 docs/CODE_HEALTH.md diff --git a/.env.example b/.env.example index 0d182b9..54728bf 100644 --- a/.env.example +++ b/.env.example @@ -31,7 +31,9 @@ COGNIFOLD_SESSION_TTL_HOURS=24 COGNIFOLD_SUPABASE_URL= COGNIFOLD_SUPABASE_KEY= COGNIFOLD_ENABLE_GRAPH_SYNC=false -COGNIFOLD_SESSION_BACKEND=memory +# To use Supabase, also set COGNIFOLD_SESSION_BACKEND=supabase above +# (do not define COGNIFOLD_SESSION_BACKEND twice in this file — the later +# value wins when sourced). # --- Cloud Run Overrides --- # These are set automatically by the CD workflow for Cloud Run deploys. @@ -40,3 +42,28 @@ COGNIFOLD_SESSION_BACKEND=memory # COGNIFOLD_TIMEOUT=300 # COGNIFOLD_SESSION_BACKEND=redis # COGNIFOLD_REDIS_URL=redis://MEMORYSTORE_IP:6379/0 + +# --- Benchmarks (optional; see docs/BENCHMARK.md "Environment variables") --- +# Reader/build model for scripts/reproduce.sh (provider:model syntax). +# Paper stack is openai:gpt-4o-mini; reproduce.sh defaults to openai:gpt-5. +# MODEL=openai:gpt-4o-mini +# Per-role overrides for scripts/parallel_longmemeval.sh: +# READER_MODEL=openai:gpt-5 +# WRITER_MODEL=openai:gpt-5 +# JUDGE_MODEL=openai:gpt-4o # judge lock — do not substitute +# RERANK_MODEL=openai:gpt-5 +# EMBED_MODEL=openai:text-embedding-3-large +# Separate endpoints/keys (optional): +# EMBEDDING_API_KEY= +# EMBEDDING_BASE_URL= +# JUDGE_API_KEY= +# JUDGE_BASE_URL= +# OpenRouter shim (exported as OPENAI_API_KEY + OpenRouter base URL by the +# benchmark shell scripts): +# OPENROUTER_API_KEY= +# Download mirrors (e.g. in China): +# HF_ENDPOINT=https://hf-mirror.com +# GITHUB_MIRROR= +# Ablation switches (read by benchmarks/shared/base_runner.py): +# COGNIFOLD_ABLATE_KNN=1 +# COGNIFOLD_ABLATE_MERGE=1 diff --git a/CLAUDE.md b/CLAUDE.md new file mode 100644 index 0000000..646342c --- /dev/null +++ b/CLAUDE.md @@ -0,0 +1,152 @@ +# CLAUDE.md + +Guidance for Claude Code (and human contributors) working in this repository. +CogniFold is a brain-inspired always-on agent memory: it folds an event stream +into a typed concept graph (`event` → `concept` → `intent`) and surfaces +proactive context without being asked. Deeper reading: `README.md` (product +overview), `docs/ARCHITECTURE.md` (technical authority), `docs/BENCHMARK.md` +(benchmark system), `docs/PHILOSOPHY.md` (design intent). + +## Commands + +```bash +uv sync # fastest install (uses uv.lock); or: pip install -e ".[dev,agent,service]" +make test # python -m pytest tests/ -v +make lint # ruff check src/ tests/ +make format # ruff format src/ tests/ +make typecheck # pyright src/ (pyright STRICT, not mypy) +make check # lint + format --check (CI runs check + typecheck + test) +make serve # ./scripts/start_server.sh (HTTP service on :8000) +bash scripts/reproduce.sh # paper benchmarks; see docs/BENCHMARK.md +bash scripts/pre-commit.sh # local quality gate (cp to .git/hooks/pre-commit) +``` + +- **Tool versions are pinned** in the `dev` extra: `ruff==0.15.17`, `pyright==1.1.408`. + Do not upgrade casually — `ruff format --check` will churn across the tree. +- Lint/format/typecheck scope is **`src/` and `tests/` only**. `benchmarks/`, + `scripts/`, and root-level `.py` files are not checked by CI. +- Python: CI, Docker, and pyright all use **3.11**. (`requires-python` says + ">=3.10" and ruff targets py39 — treat 3.11 as the real floor.) + +## Architecture map + +Write path: `Generator/Importer → Event stream → CognifoldAgent (LangGraph) → +UpdatePlan → PlanExecutor (validate → atomic execute → rollback) → ConceptGraph`. +Read path: `query → MemoryQueryAgent (legacy|bm25|semantic|hybrid|agentic) → +GraphTraverser → scorer → ContextAssembler → QueryResult`. Proactive read: +`HierarchicalContextSelector` bands (immediate/working/background) — no query needed. + +Packages under `src/cognifold/` (~40K lines, 150 files): + +| Package | Role | +|---|---| +| `models/` | Pydantic schemas: Node, Edge, Event, UpdatePlan (bottom layer, no internal deps) | +| `graph/` | NetworkX MultiDiGraph wrapper, persistence, validation, metrics | +| `scoring/` | PageRank + recency ranking, context window selection | +| `executor/` | Plan validation + atomic execution with rollback | +| `agent/` | LangGraph write-path agent, prompts, domain configs | +| `query/` | MemoryQueryAgent, strategies, assembly, LLM calls | +| `retrieval/` | BM25, hybrid RRF fusion, agentic multi-round | +| `embeddings/` | Provider ABC: Gemini/OpenAI/Mock, optional FAISS | +| `pipeline/` | Orchestration — `classic.Pipeline` (public API, MCP) and `LayeredPipeline` (CLI run, service) are BOTH alive by design | +| `intent/` | Intent-to-action queue, executor, calibrator | +| `symbolic/` | Symbolic belief/state trackers (used by benchmarks + query) | +| `temporal/` | Date parsing, temporal entity extraction | +| `service/` | FastAPI HTTP service: routes, session stores (file/redis/supabase), SSE | +| `mcp/` | MCP server (`cognifold-mcp`): remember/query/graph_stats/list_intents | +| `cli/` | `cognifold` CLI: run, query, generate, replay, serve, client | +| `generator/` `importers/` `simulator/` `replay/` | Event generation, data import, timeline sim, HTML replay | +| `brain/` | `memory_coverage.json` — **single source of truth** for the coverage figure (pages.yml copies it to docs-site) | + +**Dependency layering rule** (`docs/ARCHITECTURE.md` §Module Dependencies): +`models → graph → scoring/executor → agent → query`; `pipeline` may depend on +everything; `service`/`cli`/`simulator` sit on top. **Nothing below `service/` +should import from `service/`** — the existing `service.llm_keys` imports in +`config.py`, `agent/`, `query/`, `utils/`, `embeddings/` are a known violation +slated for cleanup (`docs/CODE_HEALTH.md` #4); do not add new ones. + +**ADR-001**: `EdgeInferenceEngine` is library code only — never auto-invoke it +during ingestion; orphan reconnection is handled by +`PlanExecutor._detect_orphan_nodes`. + +## Conventions + +- Every new file starts with `from __future__ import annotations` (128/150 + files already do). PEP 604 unions (`str | None`), builtin generics + (`dict[str, Any]`), `if TYPE_CHECKING:` for cycle-avoiding imports. +- Pyright runs **strict** on `src/` — annotate everything, including `-> None`. +- Google-style docstrings (`Args:` / `Returns:` / `Raises:`) on modules, + classes, and public functions. +- Logging: `logger = logging.getLogger(__name__)` at module level; call + `cognifold.logging.setup_logging()` once at process entry. Never use + structlog directly — the stdlib integration handles JSON output. +- **Pydantic only in `models/`** (wire/graph data). Configuration and internal + value objects use stdlib `@dataclass` (48 files follow this). +- Optional heavy deps (structlog, redis, supabase, faiss, langgraph) are + imported lazily with graceful fallback (e.g. hybrid retrieval degrades to + BM25 without an embedder). Keep that property. +- Node IDs use typed prefixes: `e-` event, `c-` concept, `i-` intent, `t-` + time; canonical form is `-` (e.g. `e-001`). +- Prompt changes for benchmarks go in `configs/_profile.yaml`, not in + Python code. + +## Testing reality + +`pytest tests/` is a **smoke suite only** (3 assertions). A green pytest does +not verify behavior. Real verification paths: + +- `PYTHONPATH=src python test_benchmarks.py` — all runners import + unified CLI +- `python scripts/test_retrieval_api.py` — all retrieval modes +- `python scripts/e2e_supabase_test.py` — service end-to-end +- Benchmark runs with `--limit 1` (cheap) before full runs (API cost) +- CI's docker-build job boots the service and polls `/health` + +When you touch a deterministic core (`graph/`, `executor/`, `scoring/`, +`retrieval/bm25`, `query/text_utils`, `config.py`), add real unit tests under +`tests/` — growing that net is an explicit goal (`docs/CODE_HEALTH.md` #1). + +## Benchmarks quick reference + +Full guide: `docs/BENCHMARK.md`. Entry points: + +| What | How | +|---|---| +| Any paper benchmark | `bash scripts/reproduce.sh {cogeval,locomo,musique,narrativeqa,tomi,babilong,mutual,streamingqa,all}` | +| LongMemEval (headline) | `bash scripts/parallel_longmemeval.sh` or the Claude Code skills `.claude/skills/longmemeval-run` / `longmemeval-iterate` | +| Datasets | run `benchmarks//download_data.py` first — runners do NOT auto-download | + +Hard rules (from the skills — keep them): +- **Judge lock**: LongMemEval/LoCoMo judge is `openai:gpt-4o` — never substitute + (breaks comparability with published baselines). +- **Branch lock**: the LongMemEval iteration campaign commits only to + `longmemeval-iter`. +- Reproducing paper numbers requires the paper stack (build `gpt-4o-mini`); + `reproduce.sh` defaults to `MODEL=openai:gpt-5` — override `MODEL` to match. + +## Branches, PRs, deployment + +- `main` and `cognifold-dev` run CI (ruff + pyright + pytest + docker health). +- Pushes to `cognifold-dev` auto-open a promote PR to `cognifold-stable`; + `cognifold-stable` deploys to Cloud Run via `cd.yml`. +- `pages.yml` deploys `docs-site/` on pushes to `main`; the brain-coverage + figure comes only from `src/cognifold/brain/memory_coverage.json`. +- PRs: single purpose, green CI, English titles/descriptions/review comments. +- PR review runs automatically (`claude-code-review.yml`); `@claude` in a + comment summons interactive help (`claude.yml`). + +## Gotchas + +- Code comments referencing `my_prompt.md §1.2` etc. now resolve to + `.claude/skills/longmemeval-iterate/references/model-config.md`. +- Default model names disagree across `config.py` / `config.example.yaml` / + `configs/cognifold.yaml` — pass models explicitly rather than trusting a + default (`docs/CODE_HEALTH.md` #13). +- `configs/` is **excluded from the Docker image** (`.dockerignore`) — the + service must not require it at runtime. +- `benchmarks/*/output/` is gitignored **except** `benchmark_results.json` and + `wrong_cases.json` (allowlisted). +- Benchmark runners must run from the repo root (relative `output_dir`). +- Undocumented ablation switches: `COGNIFOLD_ABLATE_KNN=1`, + `COGNIFOLD_ABLATE_MERGE=1` (read in `benchmarks/shared/base_runner.py`). +- `cognifold.utils.embeddings` is deprecated (import-time warning) — new code + uses `cognifold.embeddings.create_provider()`. diff --git a/docs/BENCHMARK.md b/docs/BENCHMARK.md index 524420d..f09b5a3 100644 --- a/docs/BENCHMARK.md +++ b/docs/BENCHMARK.md @@ -2,9 +2,9 @@ **This document is an extension of `CLAUDE.md`. Read `CLAUDE.md` first.** -The benchmark system evaluates Cognifold's memory against established datasets. This file is the **entry point** — it tells you what benchmarks exist, where to find everything, and what to update when making changes. - -Detailed documentation lives in `docs/benchmark/`. +The benchmark system evaluates Cognifold's memory against established datasets. +This file is the **entry point** — what benchmarks exist, how to run them, what +to update when making changes, and how to contribute a new one. --- @@ -12,7 +12,7 @@ Detailed documentation lives in `docs/benchmark/`. Headline numbers are as reported in the technical report — [arXiv:2605.13438v3](https://arxiv.org/abs/2605.13438), *CogniFold: Always-On Proactive Memory via Cognitive Folding*. The paper is the source of truth; the Implementation Status table below is the internal tracker and is kept consistent with it. -> **Why these numbers (not the highest we can get).** The reported configuration is the one that preserves proactive **intent/intention generation** end-to-end, not the per-benchmark maximum. Several older benchmarks — ToMi in particular — are easy to drive much higher with a task-specialized reader, but that path encourages auto-loop hallucination (the reader confabulates to satisfy the metric instead of reading memory). We report the proactive-substrate stack so the numbers reflect the always-on memory thesis rather than a benchmark-tuned ceiling. See PR discussion for detail. +> **Why these numbers (not the highest we can get).** The reported configuration is the one that preserves proactive **intent/intention generation** end-to-end, not the per-benchmark maximum. Several older benchmarks — ToMi in particular — are easy to drive much higher with a task-specialized reader, but that path encourages auto-loop hallucination (the reader confabulates to satisfy the metric instead of reading memory). We report the proactive-substrate stack so the numbers reflect the always-on memory thesis rather than a benchmark-tuned ceiling. ### LongMemEval (Table 5) — J-Score, N=500 @@ -50,296 +50,246 @@ Stack: `gpt-4o-mini` extraction/reader, `text-embedding-3-small`. --- -## ⚠️ ALWAYS pass `--event-stream` (every benchmark, not just LoCoMo) +## ⚠️ Consolidation and `--event-stream` — how it actually works + +Inter-session consolidation (`merge_similar_concepts` + `prune_orphan_concepts`) +is central to the always-on memory thesis, but the mechanism differs by runner: + +- **LoCoMo**: consolidation is gated behind the `--event-stream` flag, which + exists **only** in `benchmarks/locomo/run_benchmark.py`. Paper-grade LoCoMo + runs MUST pass it (`scripts/reproduce.sh locomo` does so automatically). + Sanity-check the log for `Inter-session consolidation:` lines. +- **All other `BenchmarkRunner`-based runners**: consolidation runs + unconditionally in the shared post-ingestion hook + (`benchmarks/shared/base_runner.py`, step after ingestion). There is no + `--event-stream` flag on these runners — passing it is an argparse error. + +Canonical LoCoMo (full 10-conv, Mem0 protocol): -All benchmark runners have `event_stream` default OFF, but **paper-grade runs MUST enable it** to activate per-session inter-session consolidation (`merge_similar_concepts` + `prune_orphan_concepts`). Canonical LoCoMo (full 10-conv, Mem0 protocol): ```bash PYTHONPATH=src python -u -m benchmarks.locomo.run_benchmark \ - --event-stream --model openai:gpt-4.1-mini + --event-stream --model openai:gpt-4o-mini ``` -Sanity check log for `Inter-session consolidation:` lines. Pre-2026-04-19 `--limit` default was `1` (silent conv-26-only truncation); fixed to `None` = all 10. If log shows `Loaded 1 conversations` on a full run → regression. + +Historical note: pre-2026-04-19 the LoCoMo `--limit` default was `1` (silent +conv-26-only truncation); it is now `None` = all 10. If a full run logs +`Loaded 1 conversations`, that regression is back. --- ## Implementation Status -| Benchmark | Location | Status | Accuracy (latest) | Primary Blocker / Note | -|-----------|----------|--------|-------------------|------------------------| -| **LoCoMo** | `benchmarks/locomo/` | Tested | **81.23% J-Score overall** (paper Table 4, Mem0 protocol, gpt-4o-mini read/write + gpt-4o-mini judge, `--event-stream`; Single-Hop 90.49 / Multi-Hop 67.38 / Temporal 78.50 / Open 50.00; F1 35.71) | vs ENGRAM 77.55 · MemOS 75.80 · Zep 75.14 | -| **LongMemEval** | `benchmarks/longmemeval/` | Tested | **93.0% J-Score overall** (paper Table 5, N=500, build gpt-4o-mini / answer gpt-5.4-mini / judge gpt-4o; SSA 100.0 / SSU 97.1 / KU 94.9 / SSP 93.3 / MS 91.0 / TR 88.7) | vs Mastra 94.9 · ENGRAM 71.4 · Zep 71.2; Chronos (High) 95.6. MS lever: see PR #26/#27 | -| **MSC** | `benchmarks/msc/` | Tested | N/A (excluded from Feb 21 full eval) | Agent concept extraction too passive | -| **BABILong** | `benchmarks/babilong/` | Tested | **85.0** (paper Fig. 4; proactive-substrate stack, not benchmark-tuned ceiling) | exceeds ARMT — fine-tuned (83.8) | -| **FutureX** | `benchmarks/futurex/` | Tested | N/A (no GT) | Pipeline verified, needs real MiroFlow | -| **MuTual** | `benchmarks/mutual/` | Tested | **93.2% acc** (N=500) | Near-SOTA (~97% GPT-4o zero-shot) | -| **MuSiQue-Ans** | `benchmarks/musique/` | Tested | **F1 58.7** (paper Fig. 4, N=500) | exceeds HippoRAG 2 (49.3) | -| **TimeQA** | `benchmarks/timeqa/` | Tested | 0.0% EM (Feb 21, n=20) | Temporal reasoning absent | -| **NarrativeQA** | `benchmarks/narrativeqa/` | Tested | **F1 0.720 / ROUGE-L 0.712** (Apr 8, N=500, GPT-4o) | Scoring normalization + summary detruncation | -| **QMSum** | `benchmarks/qmsum/` | Tested | F1=0.143, ROUGE-L=0.139 (N=281) | Gemini thinking-token truncation unresolved | -| **SocialIQA** | `benchmarks/socialiqa/` | Tested | **78.4% acc** (N=500) | LLM internal commonsense sufficient | -| **ToMi** | `benchmarks/tomi/` | Tested | **83.5 EM** (paper Fig. 4; proactive-substrate stack — a task-specialized reader scores far higher but invites auto-loop confabulation, so not used) | exceeds AutoToM (80.2) | -| **SafetyBench** | `benchmarks/safetybench/` | Tested | **94.3% acc** (N=35) | Exceeds GPT-4 zero-shot (88.9%); direct mode | -| **StreamingQA** | `benchmarks/streamingqa/` | Tested | **78.4% EM / F1 0.573** (N=500) | Answer-seeded fact events + containment EM | -| **RGB** | `benchmarks/rgb/` | Tested | 80.0% EM / F1 0.860 (N=20, pilot) | Wave 7 fix | -| **CogEval-Bench** (structural) | `papers/cognifold-neurips2025/` | Validated | **Harmony 0.476, Purity 0.361, Proactivity 0.614, Compression 4.6×** (6 systems × 6 scenarios, GPT-4o-mini) | Only CogniFold non-zero on Purity + Proactivity; 5-tier hierarchy revealed | - -See [docs/benchmark/results.md](benchmark/results.md) for detailed experiment results and known issues. +Nine benchmark directories exist in-tree today. Rows for retired benchmarks +(MSC, FutureX, TimeQA, QMSum, SocialIQA, SafetyBench, RGB) keep their last +recorded numbers for the paper's record but have **no code in this repo**. + +| Benchmark | Location | Status | Accuracy (latest) | Note | +|-----------|----------|--------|-------------------|------| +| **LoCoMo** | `benchmarks/locomo/` | In-tree, tested | **81.23% J-Score overall** (paper Table 4) | vs ENGRAM 77.55 · MemOS 75.80 · Zep 75.14 | +| **LongMemEval** | `benchmarks/longmemeval/` | In-tree, tested | **93.0% J-Score overall** (paper Table 5, N=500) | headline benchmark; entry point is `run_eval.py`, not `run_benchmark.py` — see below | +| **CogEval-Bench** | `benchmarks/cogeval/` | In-tree, tested | **Harmony 0.476, Purity 0.361, Proactivity 0.614, 4.6×** | dataset is generated (`generate_dataset.py`), not downloaded | +| **BABILong** | `benchmarks/babilong/` | In-tree, tested | **85.0** (paper Fig. 4) | exceeds ARMT — fine-tuned (83.8) | +| **MuTual** | `benchmarks/mutual/` | In-tree, tested | **93.2% acc** (N=500) | cleanest runner — use as the template | +| **MuSiQue-Ans** | `benchmarks/musique/` | In-tree, tested | **F1 58.7** (paper Fig. 4, N=500) | exceeds HippoRAG 2 (49.3) | +| **NarrativeQA** | `benchmarks/narrativeqa/` | In-tree, tested | **F1 0.720 / ROUGE-L 0.712** (N=500) | | +| **ToMi** | `benchmarks/tomi/` | In-tree, tested | **83.5 EM** (paper Fig. 4) | exceeds AutoToM (80.2) | +| **StreamingQA** | `benchmarks/streamingqa/` | In-tree, tested | **78.4% EM / F1 0.573** (N=500) | | +| MSC / FutureX / TimeQA / QMSum / SocialIQA / SafetyBench / RGB | *(removed)* | Retired | historical | code no longer in tree | --- ## File Map -### Documentation (docs/) - -| File | Content | -|------|---------| -| **`docs/BENCHMARK.md`** | **This file** — entry point, status overview, checklists | -| `docs/benchmark/status.md` | **Implementation progress tracker** (15/15 done, accuracy data, CLI reference) | -| `docs/benchmark/architecture.md` | System architecture, core components, profile schema, conventions | -| `docs/benchmark/dataset-catalog.md` | All 15 planned datasets across 6 categories | -| `docs/benchmark/results.md` | Experiment results, current status details, known issues | -| `docs/benchmark/phase12-log.md` | Phase 12 detailed work log (historical record) | -| **`test_benchmarks.py`** | **Test suite** — verifies all 15 runners, unified CLI args, data files | - -### Benchmark Code (benchmarks/) +### Benchmark code (`benchmarks/`) ``` benchmarks/ -├── locomo/ -│ ├── run_benchmark.py # Main runner (agent mode) -│ ├── download_data.py # Fetches locomo10.json from GitHub -│ ├── locomo10.json # Dataset (10 conversations, ~66K lines) -│ └── README.md -├── longmemeval/ -│ ├── __init__.py -│ └── run_eval.py # Runner (batch + turn modes) -├── msc/ -│ ├── run_benchmark.py # Main runner (speaker-aware) -│ ├── download_data.py # Downloads ParlAI dataset -│ ├── data/ # Downloaded data (501 conversations) -│ └── README.md -├── babilong/ -│ ├── run_benchmark.py # Runner (direct/batch/agent modes) -│ ├── download_data.py # Downloads from HuggingFace -│ ├── data/ # Downloaded data -│ └── README.md -├── futurex/ -│ ├── run_benchmark.py # Async runner (simulated tools) -│ ├── run_miroflow_benchmark.py # Real MiroFlow entry point -│ ├── miroflow_adapter.py # Cognifold <> MiroFlow adapter -│ ├── cognifold_orchestrator.py # Custom MiroFlow orchestrator -│ ├── download_data.py -│ └── futurex_data.jsonl # 96 prediction tasks -├── mutual/ -│ ├── run_benchmark.py # MC dialogue reasoning (4 choices) -│ ├── download_data.py # GitHub zip download, JSON parser -│ ├── data/ # 886 dev examples -│ └── README.md -├── musique/ -│ ├── run_benchmark.py # Multi-hop QA (EM/F1) -│ ├── download_data.py # HuggingFace download -│ ├── data/ # 4,834 examples -│ └── README.md -├── timeqa/ -│ ├── run_benchmark.py # Time-sensitive QA (EM/F1) -│ ├── download_data.py # GitHub JSONL download -│ ├── data/ # 2,997 easy + hard examples -│ └── README.md -├── narrativeqa/ -│ ├── run_benchmark.py # Long-form QA (ROUGE-L/F1) -│ ├── download_data.py # HuggingFace download -│ ├── data/ # 10,557 examples -│ └── README.md -├── qmsum/ -│ ├── run_benchmark.py # Meeting summarization (ROUGE-L) -│ ├── download_data.py # GitHub JSONL + HF fallback -│ ├── data/ # 35 meetings, 281 queries -│ └── README.md -├── socialiqa/ -│ ├── run_benchmark.py # MC social reasoning (3 choices) -│ ├── download_data.py # HF + parquet fallback -│ ├── data/ # 1,954 validation examples -│ └── README.md -├── tomi/ -│ ├── run_benchmark.py # Theory of Mind (EM) -│ ├── download_data.py # GitHub + clone/generate fallback -│ ├── data/ # 2,988 generated examples -│ └── README.md -├── safetybench/ -│ ├── run_benchmark.py # Safety MC (4 choices, no GT) -│ ├── download_data.py # HuggingFace download -│ ├── data/ # 11,435 test examples -│ └── README.md -├── streamingqa/ -│ ├── run_benchmark.py # Temporal knowledge QA -│ ├── download_data.py # GCS download (requires manual) -│ └── README.md -└── rgb/ - ├── run_benchmark.py # Noise robustness (EM/F1) - ├── download_data.py # GitHub (currently 404) - └── README.md +├── shared/ # THE shared infrastructure — read this first +│ ├── base_runner.py # BenchmarkRunner ABC + run() pipeline + LLM helpers +│ ├── baseline_runner.py # Direct-LLM / RAG baselines +│ ├── graph_evolution_tracker.py +│ └── stats_utils.py # wilson_ci, bootstrap_ci_mean +├── _utils.py # embedding resolution (the canonical embedder factory) +├── analysis_utils.py # wrong-case enrichment + save_wrong_cases +├── compare_fast.py # classic-vs-fast ingestion A/B +├── babilong/ cogeval/ locomo/ longmemeval/ +├── musique/ mutual/ narrativeqa/ +├── streamingqa/ tomi/ +└── scripts/plot_graph_evolution.py ``` -### Config Profiles (configs/) - -| File | Benchmark | -|------|-----------| -| `configs/locomo_profile.yaml` | LoCoMo | -| `configs/longmemeval_profile.yaml` | LongMemEval | -| `configs/msc_profile.yaml` | MSC | -| `configs/babilong_profile.yaml` | BABILong | -| `configs/futurex_profile.yaml` | FutureX | -| `configs/mutual_profile.yaml` | MuTual | -| `configs/musique_profile.yaml` | MuSiQue-Ans | -| `configs/timeqa_profile.yaml` | TimeQA | -| `configs/narrativeqa_profile.yaml` | NarrativeQA | -| `configs/qmsum_profile.yaml` | QMSum | -| `configs/socialiqa_profile.yaml` | SocialIQA | -| `configs/tomi_profile.yaml` | ToMi | -| `configs/safetybench_profile.yaml` | SafetyBench | -| `configs/streamingqa_profile.yaml` | StreamingQA | -| `configs/rgb_profile.yaml` | RGB | - -### Core Modules (src/cognifold/) - -These are **shared** by all benchmarks. Changes here affect everything. +Each benchmark dir: `run_benchmark.py` (runner), `download_data.py` (dataset +fetch), `README.md`. Exceptions: **longmemeval** uses `run_eval.py` + +`symbolic_resolver.py` + `runs/