Skip to content

fix: share a single BM25 model between indexing and search - #127

Merged
GoodbyePlanet merged 2 commits into
mainfrom
fix/shared-bm25-instance
Sep 13, 2026
Merged

GoodbyePlanet merged 2 commits into
mainfrom
fix/shared-bm25-instance

Conversation

@GoodbyePlanet

Copy link
Copy Markdown
Owner

Closes #70

Problem

Two Bm25("Qdrant/bm25") models were loaded per process:

  • server/main.py constructed a BM25SparseProvider() in the lifespan and stored it in server/state.py, where the search tools picked it up via get_sparse_provider().
  • IndexPipeline.__init__ resolved the separate module singleton in server/embeddings/bm25.py via get_sparse_embedding_provider().

On top of the doubled memory, close_sparse_embedding_provider() only cleared the module singleton — the copy held in state.py lived for the whole process.

Change

Drop the state.py holder entirely and route both callers through get_sparse_embedding_provider(), mirroring how the dense provider is already shared (no state.py entry; every caller goes through the factory singleton). The lifespan now just calls the getter to warm the model at startup, so the first query still doesn't pay the load cost, and teardown releases the only instance.

Sharing one model across indexing and search needs no extra locking: passage_embed / query_embed are pure over the loaded vocabulary and already run in an executor thread.

  • server/state.py — removed _sparse_provider, get_sparse_provider(), set_sparse_provider()
  • server/tools/search.py — both call sites use the shared getter
  • server/main.py — warm the singleton instead of building a second one
  • tests/tools/test_search.py — patch targets repointed (mechanical)
  • tests/embeddings/test_bm25_singleton.py — new guard test: stubs Bm25 so nothing is downloaded, asserts the pipeline and search modules resolve to the same instance, plus singleton and close behaviour

A second commit syncs uv.lock's project version to the 1.3.1 release, which release-please left at 1.2.2.

Testing

uv run pytest — 365 passed. Ruff is not installed in this environment, so no lint run.

🤖 Generated with Claude Code

GoodbyePlanet and others added 2 commits September 12, 2026 13:19
The server lifespan constructed its own BM25SparseProvider and stored it in
server/state.py for the search tools, while IndexPipeline resolved the separate
module singleton in server/embeddings/bm25.py. Two Bm25("Qdrant/bm25") models
were loaded per process, and close_sparse_embedding_provider() only released
the module one.

Drop the state.py holder and route both callers through
get_sparse_embedding_provider(), mirroring how the dense provider is already
shared. The lifespan now just warms that singleton so the first query still
does not pay the model-load cost.

Sharing one model across indexing and search is safe: passage_embed and
query_embed are pure over the loaded vocabulary and already run in an executor
thread, so no additional locking is needed.

Closes #70

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
release-please bumps pyproject.toml only; the lockfile's own project version
entry was still 1.2.2.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@GoodbyePlanet
GoodbyePlanet merged commit a41cf7f into main Sep 13, 2026
2 checks passed
@GoodbyePlanet
GoodbyePlanet deleted the fix/shared-bm25-instance branch September 13, 2026 11:33
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Two separate BM25 model instances are loaded (double memory)

1 participant