Skip to content
cair-vinuniPublic

About

Repository associated with paper titled "RoboGaze: Evaluating Robot World Models via Structured Vision-Language Analysis"

Topics

Resources

Stars

151 stars

Watchers

1 watching

Forks

Latest commit

 

History

5 Commits

Folders and files

Repository files navigation

RoboGaze: Evaluating Robot World Models via Structured Vision-Language Analysis

RoboGaze overview

Project Page arXiv Code Released

RoboGaze is a training-free, multi-agent VLM framework for diagnosing failures in generated robot-manipulation videos. Given a task instruction, an initial frame, and a generated execution video, RoboGaze produces a structured report that explains what failed, when it failed, why it failed, and how severe the failure is.

This repository contains the official RoboGaze code release, including the local inference pipeline, OpenAI-compatible VLM client, batch runners, evaluation utilities, and static project page assets.

Overview

Robot world models can generate visually plausible manipulation rollouts that still violate task logic, physical plausibility, robot-body consistency, or object-scene consistency. Scalar metrics and monolithic VLM judges often miss these errors or over-report failures in clean clips.

RoboGaze addresses this with a three-stage diagnostic pipeline:

  1. Task-scene grounding: parses the instruction and initial frame into task memory, scene memory, expected subgoals, visible objects, robot parts, and layout information.
  2. Specialist routing: identifies suspicious temporal spans and dispatches them to dimension-specific agents over a robotics failure taxonomy.
  3. Critic verification: re-examines candidate glitches, rejects weak hypotheses, merges duplicates, refines temporal boundaries, and emits a final structured report.

RoboGaze three-stage pipeline

The released implementation is intentionally transparent: intermediate JSON files and generated view clips are written to disk so that failed or ambiguous runs can be inspected.

Table of Contents

Installation

1. Clone and enter the repository

git clone https://github.com/cair-vinuni/RoboGaze.git
cd RoboGaze

2. Install core dependencies

RoboGaze uses uv for reproducible Python environment management.

uv sync

To serve the local VLM from the same environment, install the optional serving dependencies:

uv sync --extra serve

3. Install system media tools

RoboGaze expects ffmpeg and ffprobe on the system path. They are used to extract video windows, frame strips, refined clips, and media metadata.

ffmpeg -version
ffprobe -version

Serve a Local VLM

The pipeline talks to an OpenAI-compatible chat-completions endpoint. The provided script starts a local vLLM server and exposes the model as gemma4 at http://localhost:8000/v1.

bash scripts/serve_vlm.sh

Run Single-Video Inference

Set the endpoint and model name, then run robogaze on one video:

export LOCAL_BASE_URL=http://localhost:8000/v1
export LOCAL_MODEL=gemma4

uv run robogaze \
  --task-instruction "Use the right hand to pick up the red cup and place it on the plate." \
  --initial-frame /path/to/initial_frame.jpg \
  --video /path/to/execution.mp4 \
  --output-dir outputs/robogaze \
  --video-id example_001 \
  --vlm-concurrency 8

You can also use the example wrapper:

bash scripts/run_example.sh

Run Batch Inference

For prepared benchmark folders, use:

bash scripts/run_all_datasets.sh

Useful batch options:

uv run python scripts/run_robogaze_datasets.py \
  --input-root inputs/robogaze_dataset \
  --datasets gr1_real gr1_sim droid_mv \
  --output-dir outputs/robogaze_dataset \
  --cache-dir cache/robogaze_dataset \
  --vlm-concurrency 8 \
  --limit 10

The batch runner writes a JSONL summary and continues after per-sample failures unless --fail-fast is passed.

Evaluate Predictions

The evaluation script compares RoboGaze reports against canonical ground-truth JSON files. It computes per-clip metrics, per-dimension metrics, temporal IoU, text-similarity matching, coverage, and clean-clip detection summaries.

export GEMINI_API_KEY=<your_key>

uv run python scripts/evaluate.py \
  --pred-dir outputs/robogaze_dataset \
  --gt-dir ground_truth \
  --out-dir report/robogaze_eval \
  --model gemini-3.1-flash-lite

Citation

If you find RoboGaze useful for your research, please cite:

@article{nguyen2026robogaze,
  title={RoboGaze: Evaluating Robot World Models via Structured Vision-Language Analysis},
  author={Nguyen, Minh-Loi and Diep, Nghiem Tuong and Nguyen, Hung Khang and Le, Minh and Thien, Doanh Le and Tran, Hoang H and Le, Dung D and Duong, Vu N and Sonntag, Daniel and Le, An Thai and others},
  journal={arXiv preprint arXiv:2606.28385},
  year={2026}
}

About

Repository associated with paper titled "RoboGaze: Evaluating Robot World Models via Structured Vision-Language Analysis"

Topics

Resources

Stars

151 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages