Skip to content

Repository files navigation

CROSS: Change-Robust Online Topological Memory for Long-Term Relocalization and Semantic Navigation

NeurIPS 2026

Jiaming Wang, Jizhuo Chen, Diwen Liu, Atharva Ghotavadekar, Jiaxuan Da, Linh Kästner, Harold Soh

National University of Singapore

Project Page arXiv NeurIPS 2026 License: MIT

A quadruped relocalizes in a crowded canteen with a map built the evening before, then navigates to a language goal

Pose-aware topological mapping for RGB-D, stereo and monocular inputs.

CROSS builds probabilistic topological maps from RGB-D, stereo or monocular camera streams. It maintains a Gaussian mixture belief over SE(3) poses, tracks multiple hypotheses, detects loop closures, and optimizes pose graphs — enabling robust long-term navigation in indoor environments.

Three sensor modes share one back end (belief, hypotheses, verified loop closure, pose graph, retrieval). Each runs with the dataset's odometry (--odometry external) or with DPVO visual odometry (--odometry visual):

mode relative pose estimator input install / run
RGB-D (default) XFeat + LightGlue + PnP-RANSAC on keyframe depth RGB-D (or RGB + predicted depth) + odometry bash install.sh, python run.py <seq>
stereo feed-forward multi-view model (VGGT-Omega, optionally Depth Anything 3); one forward pass registers the current view against all retrieved keyframes, the known stereo baseline fixes the metric scale stereo pairs (or monocular + odometry) bash install.sh --stereo, python run.py <seq> --mode stereo
mono metric two-view matching on learned (Depth Anything 3) keyframe depth, feed-forward fallback colour images only bash install.sh --mono, python run.py <seq> --mode mono

New here? Start with the usage guide: modes, configs, multi-session mapping, and running the benchmark.

The stereo mode tolerates lighting, weather and viewpoint changes that break keypoint matching; it was developed as CROSS-stereo and is merged here (see Stereo mode). The mono mode and visual odometry come from CROSS-mono (see Mono mode and visual odometry). A fixed benchmark of all modes and baselines (mapping accuracy, multi-session localization, relocalization success) is in benchmark/; it is still being extended.

Key Features

  • Multi-hypothesis tracking — Gaussian mixture model (GMM) over SE(3) with evidence-driven lifecycle (birth, realization, removal).
  • Loop closure — Overlap-based detection with asynchronous pose graph optimization (GTSAM) and hypothesis merging.
  • Topological planning — Lightweight graph over keyframes with odometry and proximity edges; supports A* and Dijkstra path planning.
  • Visual place recognition — Keyframe database with embedding-based retrieval for relocalization.
  • Semantic memory — Text-conditioned object search across the map using open-vocabulary detectors.
  • Verified loop closure — prior, in-pass and posterior consistency tests at one chi-square level, with a noise model calibrated without ground truth from about a minute of the robot's own data.
  • Stereo and mono modes — learned multi-view relative poses with stereo scale anchors; colour-only operation with learned metric depth.
  • Visual odometry — optional DPVO motion source for every mode, so no odometry input is needed.
  • Benchmark — one protocol (T1 / T2 / T3) over KITTI, OpenLORIS-Scene, ROVER and SimChange, with ORB-SLAM3, RTAB-Map, MASt3R-SLAM and VGGT-SLAM 2.0 baselines.
  • Multiple dataset formats — R3D, ROS bags, OpenLORIS, TUM RGB-D, posed RGB-D folders; stereo: KITTI raw, TartanAir V2, Virtual KITTI 2, SimChange.

Architecture

cross/
├── core/           # System pipeline, hypothesis management, PGO, planning
├── cv/             # Pose estimation (PnP; stereo mode: feed-forward + stereo scale), feature extraction, detection
├── pipeline.py     # Sensor mode x odometry source: builds a session around the back end
├── mono/           # Mono mode: DPVO visual odometry, DA3 depth, two-view relocalization, real-time runner
├── db/             # Keyframe database and visual place recognition
├── dataloader/     # Dataset loaders (R3D, ROS bag, OpenLORIS, TUM, posed RGB-D; stereo sequences)
├── utils/          # Math (Lie algebra, rotations), profiling, camera models
└── visualization/  # Rerun-based 3D visualization, graph plotting
benchmark/          # Evaluation protocol, dataset converters, system runners, results page
configs/            # Default, stereo, outdoor, noise and mono-profile configs
docs/               # Usage guide
scripts/            # Map-and-relocalize harnesses, dataset converters, baselines, figures

Installation

Quick Start (recommended)

git clone https://github.com/jiaming-ai/CROSS.git
cd CROSS
bash install.sh

The install script handles everything: creates a virtual environment via uv, installs PyTorch with CUDA, installs CROSS and its dependencies, and builds GTSAM.

For the stereo mode add --stereo (installs the vendored VGGT-Omega and downloads its weights to models/VGGT-Omega/; the checkpoint is gated: request access at facebook/VGGT-Omega and huggingface-cli login first) and optionally --da3 for the Depth Anything 3 backend:

bash install.sh --stereo          # RGB-D + stereo mode
bash install.sh --stereo --da3    # + Depth Anything 3 backend
bash install.sh --mono            # + mono mode and visual odometry (builds DPVO; needs nvcc matching PyTorch's CUDA)
Install options
# CPU-only (no CUDA)
bash install.sh --cpu

# Specific CUDA version, it should match the cuda version on your PC, i.e. nvcc --version
CUDA_VERSION=cu121 bash install.sh

Manual Installation

# 1. Create and activate environment
uv venv --python 3.11
source .venv/bin/activate

# 2. Install PyTorch (match your CUDA version — check with nvcc --version)
uv pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu124

# 3. Install CROSS
uv pip install -e ".[all]"

# 4. Install GTSAM (required for pose graph optimization)
git clone --depth 1 https://github.com/borglab/gtsam.git thirdparty/gtsam
cd thirdparty/gtsam
uv pip install -r python/dev_requirements.txt
mkdir -p build && cd build
cmake .. -DGTSAM_BUILD_PYTHON=1 -DGTSAM_PYTHON_VERSION=3.11 -GNinja
ninja python-install
cd ../../..

Optional Dependencies

# Object detection (semantic memory)
uv pip install -e ".[detection]"

# R3D recording support
uv pip install -e ".[recording]"

Usage

Run mapping on a dataset

uv run python run.py data/r3d/lab2.r3d

Options:

Flag Description
--no-viz Disable Rerun visualization
--frames N Process only the first N frames
--start N Start from frame N
--mode {rgbd,stereo,mono} Sensor mode: PnP on depth (default), the stereo mode (layers configs/stereo.yaml), or colour only
--odometry {external,visual} Motion source: the dataset's odometry (default) or DPVO visual odometry
--config A.yaml B.yaml Config layers on top of the defaults, merged left to right
--loader {r3d,rosbag,loris,tum,posed,stereo} Force dataset loader (default: auto-detect)
--baseline B Stereo mode, SimChange sequences: which rendered stereo baseline to use
--snr FLOAT Signal-to-noise ratio for R3D datasets
--async Enable async step pipeline

Interactive demo

python examples/demo.py

Provides a REPL with commands: l (load sequence), s (set start), g (go/run), p (plan path), v (visualize graph), q (quit).

Multi-session mapping

uv run python examples/multi_session.py

Save, load, and plan

uv run python examples/planner.py --map-scene data/r3d/lab_obj.r3d --reloc-scene data/rosbag/lab_office_dog

Benchmark: mapping accuracy and relocalization success

Two harnesses build a map on one sequence and measure relocalization success on another (CROSS protocol: independent 100-frame trials that start without knowing the pose; a trial succeeds when its final estimate is within --r-d of the pose the map implies, 2 m indoors), plus the keyframe ATE of the map:

  • scripts/map_and_reloc_rgbd.py: RGB-D mode on posed RGB-D folders (rgb/, depth/, poses_left.txt, optional odom_left.txt with the robot's own odometry, calib.json);
  • scripts/map_and_reloc.py: stereo sequences (SimChange, KITTI raw, TartanAir V2, Virtual KITTI 2) or posed RGB-D folders, with --estimator ff (stereo mode) or --estimator pnp --pnp-depth gt|sgbm.

Both take --mode mono (colour / left image only) and --odometry visual (no odometry input, see Mono mode and visual odometry). The full benchmark protocol (T1 mapping ATE, T2 multi-session localization, T3 relocalization success; KITTI, OpenLORIS-Scene, ROVER, SimChange; baselines) is in benchmark/.

# OpenLORIS-Scene (package format) -> posed RGB-D folders; the robot's wheel odometry is kept and used as odometry
python scripts/datasets/convert_openloris.py data/openloris/home1-1 data/posed/home1-1
python scripts/datasets/convert_openloris.py data/openloris/home1-2 data/posed/home1-2
python scripts/map_and_reloc_rgbd.py --map data/posed/home1-1 --query data/posed/home1-2 --out outputs/home
# the same in the stereo mode, monocular (metric scale from the wheel odometry and map keyframe pairs)
python scripts/map_and_reloc.py --map data/posed/home1-1 --query data/posed/home1-2 --out outputs/home_ff --estimator ff \
    --obs-min-translation 0.3 --obs-min-rotation 0.15 --obs-max-interval 3 \
    --set pose_est.ff.use_odom_anchor=true pose_est.ff.use_map_anchors=true pose_est.ff.n_ref_anchors=0 pose_est.ff.use_curr_anchor=false
# stereo mode on a SimChange scene (0.3 m baseline)
python scripts/map_and_reloc.py --map data/sim/hssd_house/map --query data/sim/hssd_house/light_night --out outputs/house_ff \
    --estimator ff --baseline 0.3 --snr 10 --obs-min-translation 0.3 --obs-min-rotation 0.15 --obs-max-interval 3 \
    --trial-len 100 --trial-stride 50 --noise-config configs/noise/hssd_house_600.yaml
# TUM RGB-D: scripts/datasets/convert_tum.py (simulated noisy odometry: --snr 10 --seed 0)

Noise calibration for a new robot (optional)

The verified loop closure uses a noise model of the relative-pose estimator and the odometry. The defaults work for wheeled indoor robots; for another platform, record about a minute of data and calibrate without ground truth:

python scripts/map_and_reloc_rgbd.py --map <seq> --query <seq> --out outputs/calib --map-end 600 --skip-reloc --dump-graph
python scripts/lc/calibrate_noise.py --graph outputs/calib/graph_s0.json --out configs/noise/my_robot.yaml
# stereo mode: record with scripts/viz/record_trace.py (graph + pass-internal poses) and pass --trace as well
# then: mapping.loop_closure.noise_file: configs/noise/my_robot.yaml

Main defaults: verified loop closure (consistency tests at one chi-square level), hypothesis 0 updated only by measurements that are more informative than the odometry chain, keyframe images stored as uint8. Observation gating (pose_est.obs_min_translation: 0.3, obs_min_rotation: 0.15, obs_max_interval_steps: 3) is on in the stereo preset (the feed-forward observation costs ~0.3 s); in the RGB-D mode it runs 1.4-1.8x faster but is off by default because it cost relocalization success on one real-robot scene (OpenLORIS home).

Stereo mode

The classical relative pose estimator (XFeat + LightGlue + PnP-RANSAC on keyframe depth) is replaced by a feed-forward multi-view geometry model (VGGT-Omega or Depth Anything 3). All retrieved reference keyframes and the current frame are processed in one forward pass so that they share one similarity gauge; the known stereo baseline of the current frame (and of stored keyframe right images) fixes the metric scale of that gauge (robust log-space estimation, cross/cv/stereo_scale.py), and a covisibility score from the predicted depth maps replaces the PnP inlier count as the measurement confidence. Because the model is heavier than PnP, the observation (retrieval + forward pass) runs at a motion-gated cadence while odometry propagates the multi-hypothesis belief in between. Without a right camera the scale comes from the odometry (previous observed frame) and from pairs of map keyframes.

python run.py data/kitti_raw/2011_09_30/2011_09_30_drive_0027_sync --mode stereo --config configs/outdoor.yaml
python run.py data/sim/lonemonk/map --mode stereo --baseline 0.3

Programmatic use differs from the RGB-D mode only by the stereo calibration and the right image:

cfg = load_config("configs/stereo.yaml")          # pose_est.type = ff (+ observation gating)
system = System(camera=camera, config=cfg, T_right_in_left=dataset.T_right_in_left)
system.step(obs={"rgb": left, "rgb_right": right, "depth": None, "conf": None,
                 "delta_pose": odom_delta, "timestamp": t})

Module-level evaluation (relative pose accuracy vs. ground truth) and the report experiments:

python scripts/eval_relpose.py --ref data/vkitti2/Scene01/clone --query data/vkitti2/Scene01/sunset \
    --estimator ff --backend vggt_omega --gaps 0,5,10,20 --n-queries 40 --out outputs/relpose/vk01_sunset_ff.json
bash scripts/run_experiments.sh all && python scripts/summarize.py && python scripts/make_figures.py

Mono mode and visual odometry

A run is defined by two independent choices (cross/pipeline.py):

--odometry external (default) --odometry visual
--mode rgbd dataset odometry; PnP on sensor depth DPVO visual odometry, metric scale from the sensor depth
--mode stereo dataset odometry; feed-forward estimator with stereo scale DPVO on the left image, metric scale from stereo (SGBM) depth
--mode mono dataset odometry; learned metric depth (Depth Anything 3) for keyframes DPVO with a learned metric-scale prior (DA3-Metric)

Every motion source feeds the same channel of the back end, the per-frame relative pose (and its covariance) that becomes the odometry chain of the pose graph; with external odometry in the RGB-D and stereo modes the frames reach System.step unchanged. The mono mode observes only colour images: relocalization geometry comes from metric two-view matching (XFeat / LighterGlue PnP on the stored keyframe depth, SuperPoint / LightGlue two-view geometry) with a feed-forward fallback (DA3) for saved-map references, coordinate charts keep an unanchored session separate from a loaded map until it is joined, and the global observation runs every few frames (the frontend pose, carried into the map frame by the last observation, is reported in between).

python run.py data/posed/home1-1 --loader posed --odometry visual            # RGB-D, no odometry input
python run.py data/posed/home1-1 --loader posed --mode mono --odometry visual # RGB only
python scripts/map_and_reloc_rgbd.py --map data/posed/home1-1 --query data/posed/home1-2 --out outputs/home_mono \
    --mode mono --odometry visual                                             # mono profile: configs/mono_benchmark_10hz.json

The DPVO weights are read from models/dpvo.pth (or --dpvo-checkpoint, CROSS_DPVO_CHECKPOINT); DPVO itself must be importable (install.sh --mono builds it under thirdparty/DPVO). With sensor or stereo depth the DPVO scale is observed every frame and the back end keeps the mode's odometry noise model; the mono mode uses the DPVO noise model of its profile. python -m cross.mono.run is the real-time monocular runner (paced input, a separate mapping process, streaming metric depth); its recommended profile is configs/mono_streaming_dpvo_v2_20hz.json, and configs/mono_benchmark_10hz.json is the same profile for offline runs at 10 Hz.

from cross.pipeline import build_session, mono_config_from_profile
session = build_session("mono", "visual", camera, SystemConfig(),
                        mono_config=mono_config_from_profile("configs/mono_benchmark_10hz.json", "models/dpvo.pth"))
for frame in dataset.replay_data():          # dict: rgb (uint8), timestamp[, depth, rgb_right, delta_pose]
    session.process(frame)
    T_c0, T_best, weights = session.belief(to_matrix)

SimChange benchmark and baselines

The multi-traversal simulator benchmark (controlled lighting, object rearrangement, background, viewpoint and traversal changes; HSSD house and restaurant, Lone Monk, classroom) lives in its own repository, SimChange (scene assets, routes, robot-motion model, Blender renderer, rearrangement quantification). Link its renders as data/sim (ln -s $SIMCHANGE_DATA/renders data/sim). Drivers for ORB-SLAM3 (stereo), RTAB-Map (RGB-D / stereo) and MASt3R-SLAM with the same trial protocol are in scripts/baselines/.

bash scripts/run_sim_experiments.sh classroom ff:0.3 pnp pnpsgbm:0.3 mast3r
bash scripts/run_sim_experiments.sh classroom orbslam3:0.3 rtabmap
python scripts/make_sim_figures.py && python scripts/make_public_tables.py

Map-construction replay

scripts/viz/record_trace.py records how a map is built and reused (belief of every hypothesis, retrievals, keyframes, edges, loop closures) for either mode; record_baseline_trace.py does the same for the baselines; build_trace_page.py renders an interactive page and render_trace_video.py MP4s.

python scripts/viz/record_trace.py --scene lonemonk --map map --variants light_night reverse \
    --out outputs/viz/lonemonk/trace_cross_stereo --baseline 0.3 --snr 10 --no-frames
python scripts/viz/build_trace_page.py --viz-root outputs/viz --scenes lonemonk --out outputs/viz/page --standalone

Datasets

OpenLORIS

Download the package format from Hugging Face (see the dataset page) and convert sequences with scripts/datasets/convert_openloris.py (above); the legacy data/loris/ loader (--loader loris) still works.

Stereo datasets

Stereo sequences are read by cross/dataloader/stereo_loader.py: KITTI raw drives (data/kitti_raw/<date>/<drive>_sync), TartanAir V2, Virtual KITTI 2 and SimChange renders (left/, right_<baseline>/, depth/, poses_left.txt, calib.json); odometry is simulated from the ground truth (--snr, optional drift) unless a posed folder provides odom_left.txt.

R3D

Place .r3d recordings in data/r3d/.

ROS Bags

Place processed ROS bag directories in data/rosbag/.

Project Structure

Module Description
cross.core.system Main pipeline: motion prior, observation, GMM filtering, keyframe insertion
cross.core.hypothesis Multi-hypothesis GMM belief, evidence tracking, loop-closure detection
cross.core.pgo Pose graph construction and GTSAM optimization
cross.core.lc_engine Asynchronous loop-closure engine (background thread)
cross.core.simple_topo Topological planning graph with proximity edges
cross.core.planner A*/Dijkstra path planning over sparse graph
cross.core.mem Text-conditioned semantic memory search
cross.db.db Keyframe database with embedding-based VPR
cross.cv.pose_est_pnp PnP-based relative pose estimation (RGB-D mode)
cross.cv.pose_est_ff Feed-forward multi-view relative pose estimation with stereo / odometry / map scale anchors (stereo mode)
cross.cv.stereo_scale Robust metric scale from calibrated anchors (stereo mode)
cross.core.lc_verify Verified loop closure: consistency tests, calibrated noise model, odometry-chain predictor
cross.dataloader.stereo_loader Stereo sequences (KITTI raw, TartanAir V2, Virtual KITTI 2, SimChange) and posed RGB-D folders
cross.pipeline Sensor mode x odometry source: sessions of the back end fed by external odometry or DPVO visual odometry
cross.mono Mono mode: DPVO frontend with learned metric scale, DA3 models, monocular two-view / feed-forward relocalization geometry, real-time runner (python -m cross.mono.run)
cross.visualization.viz_rr Rerun-based 3D visualization

Citation

If you use CROSS in your research, please cite:

@inproceedings{wang2026cross,
  title     = {Change-Robust Online Topological Memory for Long-Term
               Relocalization and Semantic Navigation},
  author    = {Wang, Jiaming and Chen, Jizhuo and Liu, Diwen and
               Ghotavadekar, Atharva and Da, Jiaxuan and K{\"a}stner, Linh
               and Soh, Harold},
  booktitle = {Advances in Neural Information Processing Systems (NeurIPS)},
  year      = {2026},
  eprint    = {2605.02227},
  archivePrefix = {arXiv}
}

License

MIT

About

Resources

Stars

8 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages