RoboGaze is a training-free, multi-agent VLM framework for diagnosing failures in generated robot-manipulation videos. Given a task instruction, an initial frame, and a generated execution video, RoboGaze produces a structured report that explains what failed, when it failed, why it failed, and how severe the failure is.
This repository contains the official RoboGaze code release, including the local inference pipeline, OpenAI-compatible VLM client, batch runners, evaluation utilities, and static project page assets.
Robot world models can generate visually plausible manipulation rollouts that still violate task logic, physical plausibility, robot-body consistency, or object-scene consistency. Scalar metrics and monolithic VLM judges often miss these errors or over-report failures in clean clips.
RoboGaze addresses this with a three-stage diagnostic pipeline:
- Task-scene grounding: parses the instruction and initial frame into task memory, scene memory, expected subgoals, visible objects, robot parts, and layout information.
- Specialist routing: identifies suspicious temporal spans and dispatches them to dimension-specific agents over a robotics failure taxonomy.
- Critic verification: re-examines candidate glitches, rejects weak hypotheses, merges duplicates, refines temporal boundaries, and emits a final structured report.
The released implementation is intentionally transparent: intermediate JSON files and generated view clips are written to disk so that failed or ambiguous runs can be inspected.
- Installation
- Serve a Local VLM
- Run Single-Video Inference
- Run Batch Inference
- Evaluate Predictions
- Citation
git clone https://github.com/cair-vinuni/RoboGaze.git
cd RoboGazeRoboGaze uses uv for reproducible Python environment management.
uv syncTo serve the local VLM from the same environment, install the optional serving dependencies:
uv sync --extra serveRoboGaze expects ffmpeg and ffprobe on the system path. They are used to extract video windows, frame strips, refined clips, and media metadata.
ffmpeg -version
ffprobe -versionThe pipeline talks to an OpenAI-compatible chat-completions endpoint. The provided script starts a local vLLM server and exposes the model as gemma4 at http://localhost:8000/v1.
bash scripts/serve_vlm.shSet the endpoint and model name, then run robogaze on one video:
export LOCAL_BASE_URL=http://localhost:8000/v1
export LOCAL_MODEL=gemma4
uv run robogaze \
--task-instruction "Use the right hand to pick up the red cup and place it on the plate." \
--initial-frame /path/to/initial_frame.jpg \
--video /path/to/execution.mp4 \
--output-dir outputs/robogaze \
--video-id example_001 \
--vlm-concurrency 8You can also use the example wrapper:
bash scripts/run_example.shFor prepared benchmark folders, use:
bash scripts/run_all_datasets.shUseful batch options:
uv run python scripts/run_robogaze_datasets.py \
--input-root inputs/robogaze_dataset \
--datasets gr1_real gr1_sim droid_mv \
--output-dir outputs/robogaze_dataset \
--cache-dir cache/robogaze_dataset \
--vlm-concurrency 8 \
--limit 10The batch runner writes a JSONL summary and continues after per-sample failures unless --fail-fast is passed.
The evaluation script compares RoboGaze reports against canonical ground-truth JSON files. It computes per-clip metrics, per-dimension metrics, temporal IoU, text-similarity matching, coverage, and clean-clip detection summaries.
export GEMINI_API_KEY=<your_key>
uv run python scripts/evaluate.py \
--pred-dir outputs/robogaze_dataset \
--gt-dir ground_truth \
--out-dir report/robogaze_eval \
--model gemini-3.1-flash-liteIf you find RoboGaze useful for your research, please cite:
@article{nguyen2026robogaze,
title={RoboGaze: Evaluating Robot World Models via Structured Vision-Language Analysis},
author={Nguyen, Minh-Loi and Diep, Nghiem Tuong and Nguyen, Hung Khang and Le, Minh and Thien, Doanh Le and Tran, Hoang H and Le, Dung D and Duong, Vu N and Sonntag, Daniel and Le, An Thai and others},
journal={arXiv preprint arXiv:2606.28385},
year={2026}
}
