Skip to content

bug: repeated Windows host crashes; bound whole-tree snapshot memory #403

Description

@2233admin

Safety / symptom

The reporter has experienced repeated whole-host crashes during Code Intel work, including another crash during an attempted fix. Do not reproduce, build, or run the full test suite on the affected Windows host. This issue tracks a verified resource-pressure defect and a separate, unproven crash attribution.

Read-only evidence

  • Earlier Windows records include Kernel-Power 41 with BugcheckCode 26 (0x1A), parameter 1 0x3F; Microsoft documents this subtype as a pagefile inpage CRC error. A separate reboot recorded bugcheck code 0. Do not conflate the two or assign either to this process without process/dump evidence.
  • A roughly 5.27 MiB minidump exists from the earlier failure, but the current account cannot read it. Process-creation telemetry for the failure window was not available. The newly reported crash has not been inspected. No dump or machine-specific path is attached.
  • crates/code-intel-cli/src/snapshot.rs: digest_worktree reads each scoped file with fs::read and retains framed contents in records (lines 1013-1099); hash_records concatenates those records into a second canonical buffer (lines 1610-1616). Peak allocation grows with whole-tree content, not just one file. The successful native-lite path builds snapshots repeatedly; this is a static call-path count, not a measured crash-time RSS.
  • MCP get_gate_verdict calls freshness (mcp_serve/handlers.rs:50-54); committed_evidence.rs:88-120 rebuilds the current worktree snapshot per such request. The idle stdio server does not automatically perform this scan.

Upstream-owned repair and acceptance

  1. In an independent, safe host/worktree, replace whole-tree retained records plus concatenated canonical buffer with incremental hashing of the same framed bytes and ordering. Keep snapshot identity and artifact-contract parity; bound peak memory by the largest file/chunk rather than the sum of all files. Review whether repeated unchanged snapshots can be avoided without weakening freshness semantics.
  2. Add deterministic digest-parity and bounded-memory checks with representative large files. Run focused tests plus cargo test --workspace --no-fail-fast and relevant integration gates on safe CI/host, not on the affected computer. Have an upstream reviewer inspect results before release.
  3. Treat the host crash as unresolved until a trusted, authorized dump analysis or contemporaneous command/process/commit telemetry establishes causation. A 0x1A/0x3F pagefile CRC can involve storage, controller, memory, or driver faults; Code Intel I/O/memory pressure may trigger that path but is not proven to be its sole cause.

No repository or OS changes, workload reproduction, dump upload, or new crash verification were performed for this report.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't workingclaimedIssue claimed by an active session (DR-0004): read the claim comment before touching it

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions