Skip to content

Latest commit

 

History

5 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

3D HAMSTER : Bridging Planning and Control in Hierarchical Vision Language Action Models through
3D Trajectory Guidance

arXiv Project Page Code Model Demo License

Dongyoon Hwang*1, Byungkun Lee*1, Dongjin Kim*1, Hyojin Jang1, Hoiyeong Jin1, Jueun Mun2,
Minho Park1, Hojoon Lee3, Hyunseung Kim1,4, Jaegul Choo†1

1KAIST AI    2POSTECH    3Holiday Robotics    4KRAFTON AI

*Equal contribution    Corresponding author

🎉 Accepted to IROS 2026 — IEEE/RSJ International Conference on Intelligent Robots and Systems


TL;DR — 3D HAMSTER is a depth-aware VLM planner that predicts metrically grounded 3D end-effector trajectories directly from a single RGB-D observation and a language instruction. Unlike 2D planners whose pixel waypoints inherit whatever depth lies beneath them, 3D HAMSTER plans in metric 3D space, so the trajectory stays geometrically grounded and can feed straight into a point-cloud low-level policy.

3D HAMSTER architecture: a depth-augmented VLM planner produces metric 3D waypoints that unproject into the point cloud consumed by the low-level policy.

This repository provides inference-only code for the 3D HAMSTER VLM planner — loading the released checkpoint, running trajectory prediction on RGB-D inputs, and a Gradio demo. (The low-level point-cloud policy is part of the full system described in the paper and is not included here.)

Highlights

  • Input: a single RGB image + a metric depth map + a language instruction.
  • Output: a metric 3D end-effector trajectory as [u, v, depth] waypoints (pixel coordinates + metric depth in meters) plus gripper actions, in structured JSON.
  • Backbone: Qwen3-VL-8B augmented with a frozen LingBot-Depth geometry encoder (DINOv2 ViT-L/14) and a dense depth-reconstruction objective.
  • Self-contained: pip install this package, download the checkpoint, and run. The geometry-encoder code is vendored in-repo and its weights ship inside the checkpoint — no separate model download and no network access at load time.

Installation

# 1. Create and activate an environment (Python >= 3.10)
conda create -n 3d_hamster python=3.11 -y
conda activate 3d_hamster

# 2. Clone and install (pulls in the vendored LingBot-Depth encoder code)
git clone https://github.com/DAVIAN-Robotics/3D_HAMSTER.git
cd 3D_HAMSTER
pip install -e .

Optional extra — pip install -e ".[perf]" adds xformers (faster attention; matches the reference setup, falls back to an equivalent path when absent).

Model Weights

The checkpoint is hosted on the Hugging Face Hub at DAVIAN-Robotics/3D_HAMSTER — a single self-contained checkpoint (9B, bf16) that bundles the Qwen3-VL LLM, the vision encoder, the geometry merger, and the frozen LingBot-Depth encoder weights.

# Download into ./ckpt
hf download DAVIAN-Robotics/3D_HAMSTER --local-dir ckpt

Quickstart — 3D Trajectory Prediction (Python API)

By default, Hamster3DPredictor performs 3D trajectory prediction: from a single RGB-D observation and a language instruction it returns a metric 3D end-effector trajectory as [u, v, depth] waypoints with per-waypoint gripper actions.

from hamster3d.inference import Hamster3DPredictor
import numpy as np
from PIL import Image

predictor = Hamster3DPredictor("ckpt/")        # device="cuda:0", bf16 by default

rgb = Image.open("examples/sample_0_rgb.png")
depth = np.load("examples/sample_0_depth.npy")          # float32, meters, shape (H, W)
instruction = open("examples/sample_0_instruction.txt").read().strip()

result = predictor.predict(rgb, depth, instruction)     # 3D trajectory prediction

print(result["waypoints"])   # [[u, v, depth], ...]  pixel u,v (0-1000) + metric depth (m)
print(result["actions"])     # ["Close Gripper", None, ..., "Open Gripper"]
print(result["raw_output"])  # raw structured-JSON string from the model

Six ready-to-run examples ship in examples/ (sample_0sample_5), each with an RGB image, a depth .npy, an instruction, and the camera intrinsics (*_camera.json).

Input / Output specification

Format
RGB PIL.Image (any resolution; auto-resized so the longest edge is 640 px)
Depth np.ndarray, float32, shape (H, W), metric depth in meters (e.g. 0.26 – 9.99). Must be aligned to the RGB frame.
Instruction free-form English string
Output dict with waypoints ([[u, v, depth], ...] — 3D trajectory), actions (gripper action or None per waypoint), raw_output (raw structured-JSON string), and the resized rgb_resized / depth_resized arrays

predict() also accepts max_new_tokens (default 1024). The Python API is dedicated to 3D trajectory prediction; for the other task styles, use the Gradio demo below.

The Gradio demo exposes all task styles via a dropdown (each sends the matching prompt and renders the result):

Task style Output Visualization
3D Trajectory (default) point_3d waypoints [u, v, depth] + gripper 2D overlay + 3D scene path
2D Trajectory point_2d waypoints [u, v] + gripper 2D overlay
3D Pointing point_3d points [u, v, depth] numbered dots + 3D scene markers
2D Pointing point_2d points [u, v] numbered dots
2D Bounding Box bbox_2d [x1, y1, x2, y2] boxes
General VQA free-form text answer text

⚠️ Depth must be metric (meters) and aligned to the RGB image. Passing disparity, normalized, or millimeter-scaled depth will silently degrade the predicted geometry.

Gradio Demo

Try it directly in your browser on 🤗 Hugging Face Spaces — no setup required. To run locally:

CUDA_VISIBLE_DEVICES=0 python scripts/trajectory_prediction_gradio.py --autoload

Open the Examples Browser tab to run the bundled examples/ (or use Manual Inference to upload your own RGB + depth): pick a sample, set the task instruction / prompt style, and inspect the predicted 2D trajectory, the 3D trajectory, and the full conversation.

Acknowledgments & Licensing

3D HAMSTER builds on several open-source projects:

This repository is released under the Apache License 2.0 (see LICENSE). The bundled LingBot-Depth and DINOv2 components are themselves Apache-2.0; their licenses and attributions are retained and apply to the corresponding code and weights.

Citation

If you find 3D HAMSTER useful, please cite our work:

@article{hwang20263dhamster,
  title={3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance},
  author={Hwang, Dongyoon and Lee, Byungkun and Kim, Dongjin and Jang, Hyojin and Jin, Hoiyeong and Mun, Jueun and Park, Minho and Lee, Hojoon and Kim, Hyunseung and Choo, Jaegul},
  journal={arXiv preprint arXiv:2606.31329},
  year={2026}
}

About

Code for "3D HAMSTER: Enabling Robust Manipulation in Hierarchical Vision Language Action Models through 3D Trajectory Guidance" (IROS 2026)

Resources

Stars

30 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages