Skip to content

Latest commit

 

History

6 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

LVMT: Video Mask Transformer for Long-term Video Segmentation

📄 Paper

Narges Norouzi1, Niccolò Cavagnero1, Idil Esen Zulfikar2, Bastian Leibe2, Gijs Dubbelman1, Daan de Geus1

¹ Eindhoven University of Technology, ² RWTH Aachen University

Overview

Long-term memory. No speed penalty.

+4.6 AP over PMT across ViT-L/B/S on OVIS, at similar FPS.
+2.4 AP over the prior state-of-the-art, DVIS-DAQ, at over 10× its speed.

LVMT Overview

We introduce the Long-term Video Mask Transformer (LVMT), an online video segmentation model built on a frozen plain Vision Transformer (ViT). A single lightweight mask decoder handles both segmentation and temporal association, without relying on dedicated tracking modules or heavy task-specific heads.

LVMT propagates information over time with Truncated Query Propagation (TQP): a GRU updates each object query independently, so every query carries its own memory across frames, and training splits long clips into chunks that keep gradients bounded while the propagated state still spans the whole clip. This lets the memory be trained on videos long enough for objects to be occluded and re-appear, which is where per-frame baselines lose track.

Installation

If you don't have Conda installed, install Miniconda and restart your shell:

wget https://repo.anaconda.com/miniconda/Miniconda3-latest-Linux-x86_64.sh
bash Miniconda3-latest-Linux-x86_64.sh

Then create the environment, activate it, and install the dependencies:

conda create -n lvmt python==3.13.2
conda activate lvmt
pip install torch==2.9.0 torchvision==0.24.0 --index-url https://download.pytorch.org/whl/cu128
python -m pip install --no-build-isolation 'git+https://github.com/facebookresearch/detectron2.git'
pip install git+https://github.com/cocodataset/panopticapi.git
python3 -m pip install -r requirements.txt

Weights & Biases (wandb) is used for experiment logging and visualization. To enable wandb, log in to your account:

wandb login

Data preparation

Download and prepare the datasets.

Usage

Evaluation

To evaluate a pre-trained LVMT model, first prepare the datasets by following the instructions in this link and download the trained weights from DINOv2 models or DINOv3 models. Once these are set up, run:

python train_net_video.py \
  --num-gpus 1 \
  --config-file /path/to/config.yaml \
  --eval-only MODEL.WEIGHTS /path/to/weight.pth \
  MODEL.BACKBONE.TEST.WINDOW_SIZE 1 \
  OUTPUT_DIR /path/to/output

🔧 Replace /path/to/config.yaml with the path to the config file.
🔧 Replace /path/to/weight.pth with the path to the checkpoint to evaluate.
🔧 Replace /path/to/output with the path to the output folder.
🔧 Change the value of --num-gpus to the number of GPUs available to you.

For detailed instructions on running evaluation on different datasets, see Evaluation.

Benchmark

To calculate the FPS and GFLOPs, run:

# DINOv2 FPS
python benchmark.py \
  --task fps \
  --config-file    /path/to/config.yaml \
  --model-weights  /path/to/weight.pth \
  --warmup-iters 100 \
  --model-type dinov2 \
  --fused-qkv

# DINOv3 FPS
python benchmark.py \
  --task fps \
  --config-file    /path/to/config.yaml \
  --model-weights  /path/to/weight.pth \
  --warmup-iters 100 \
  --model-type dinov3 \
  --fused-qkv

# DINOv2 GFLOPs
export TIMM_FUSED_ATTN=0
python benchmark.py \
  --task flops \
  --config-file    /path/to/config.yaml \
  --model-weights  /path/to/weight.pth \
  --model-type dinov2

# DINOv3 GFLOPs
python benchmark.py \
  --task flops \
  --config-file    /path/to/config.yaml \
  --model-weights  /path/to/weight.pth \
  --model-type dinov3

🔧 Replace /path/to/config.yaml with the path to the config file.
🔧 Replace /path/to/weight.pth with the path to the checkpoint to evaluate.

Demo

We provide example visualizations below.

Upcoming Features

- [x] Inference code
- [x] Flops and FPS code
- [x] DINOv2 and DINOv3 model zoo and code
- [ ] Visualization code
- [ ] Training code

Model Zoo

We provide pre-trained weights for both DINOv2- and DINOv3-based LVMT models.

Citation

If you find this work useful in your research, please cite it using the BibTeX entry below:

@article{Norouzi2026LVMT,
  author    = {Norouzi, Narges and Cavagnero, Niccol\`{o} and Zulfikar, Idil and Leibe, Bastian and Dubbelman, Gijs and {de Geus}, Daan},
  title     = {{LVMT: Video Mask Transformer for Long-term Video Segmentation}},
  journal   = {arXiv},
  year      = {2026},
}

Acknowledgements

This project builds upon code from the following libraries and repositories:

About

LVMT: Video Mask Transformer for Long-term Video Segmentation

Resources

Stars

9 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages