Narges Norouzi1, Niccolò Cavagnero1, Idil Esen Zulfikar2, Bastian Leibe2, Gijs Dubbelman1, Daan de Geus1
¹ Eindhoven University of Technology, ² RWTH Aachen University
+4.6 AP over PMT across ViT-L/B/S on OVIS, at similar FPS.
+2.4 AP over the prior state-of-the-art, DVIS-DAQ, at over 10× its speed.
We introduce the Long-term Video Mask Transformer (LVMT), an online video segmentation model built on a frozen plain Vision Transformer (ViT). A single lightweight mask decoder handles both segmentation and temporal association, without relying on dedicated tracking modules or heavy task-specific heads.
LVMT propagates information over time with Truncated Query Propagation (TQP): a GRU updates each object query independently, so every query carries its own memory across frames, and training splits long clips into chunks that keep gradients bounded while the propagated state still spans the whole clip. This lets the memory be trained on videos long enough for objects to be occluded and re-appear, which is where per-frame baselines lose track.
If you don't have Conda installed, install Miniconda and restart your shell:
wget https://repo.anaconda.com/miniconda/Miniconda3-latest-Linux-x86_64.sh
bash Miniconda3-latest-Linux-x86_64.shThen create the environment, activate it, and install the dependencies:
conda create -n lvmt python==3.13.2
conda activate lvmt
pip install torch==2.9.0 torchvision==0.24.0 --index-url https://download.pytorch.org/whl/cu128
python -m pip install --no-build-isolation 'git+https://github.com/facebookresearch/detectron2.git'
pip install git+https://github.com/cocodataset/panopticapi.git
python3 -m pip install -r requirements.txtWeights & Biases (wandb) is used for experiment logging and visualization. To enable wandb, log in to your account:
wandb loginDownload and prepare the datasets.
To evaluate a pre-trained LVMT model, first prepare the datasets by following the instructions in this link and download the trained weights from DINOv2 models or DINOv3 models. Once these are set up, run:
python train_net_video.py \
--num-gpus 1 \
--config-file /path/to/config.yaml \
--eval-only MODEL.WEIGHTS /path/to/weight.pth \
MODEL.BACKBONE.TEST.WINDOW_SIZE 1 \
OUTPUT_DIR /path/to/output🔧 Replace /path/to/config.yaml with the path to the config file.
🔧 Replace /path/to/weight.pth with the path to the checkpoint to evaluate.
🔧 Replace /path/to/output with the path to the output folder.
🔧 Change the value of --num-gpus to the number of GPUs available to you.
For detailed instructions on running evaluation on different datasets, see Evaluation.
To calculate the FPS and GFLOPs, run:
# DINOv2 FPS
python benchmark.py \
--task fps \
--config-file /path/to/config.yaml \
--model-weights /path/to/weight.pth \
--warmup-iters 100 \
--model-type dinov2 \
--fused-qkv
# DINOv3 FPS
python benchmark.py \
--task fps \
--config-file /path/to/config.yaml \
--model-weights /path/to/weight.pth \
--warmup-iters 100 \
--model-type dinov3 \
--fused-qkv
# DINOv2 GFLOPs
export TIMM_FUSED_ATTN=0
python benchmark.py \
--task flops \
--config-file /path/to/config.yaml \
--model-weights /path/to/weight.pth \
--model-type dinov2
# DINOv3 GFLOPs
python benchmark.py \
--task flops \
--config-file /path/to/config.yaml \
--model-weights /path/to/weight.pth \
--model-type dinov3🔧 Replace /path/to/config.yaml with the path to the config file.
🔧 Replace /path/to/weight.pth with the path to the checkpoint to evaluate.
We provide example visualizations below.
- [x] Inference code
- [x] Flops and FPS code
- [x] DINOv2 and DINOv3 model zoo and code
- [ ] Visualization code
- [ ] Training code
We provide pre-trained weights for both DINOv2- and DINOv3-based LVMT models.
- DINOv2 Models - DINOv2-based models and pre-trained weights.
- DINOv3 Models - DINOv3-based models and pre-trained weights.
If you find this work useful in your research, please cite it using the BibTeX entry below:
@article{Norouzi2026LVMT,
author = {Norouzi, Narges and Cavagnero, Niccol\`{o} and Zulfikar, Idil and Leibe, Bastian and Dubbelman, Gijs and {de Geus}, Daan},
title = {{LVMT: Video Mask Transformer for Long-term Video Segmentation}},
journal = {arXiv},
year = {2026},
}This project builds upon code from the following libraries and repositories:
- EoMT (MIT License)
- VidEoMT (MIT License)
- PMT (MIT License)
- Hugging Face Transformers (Apache-2.0 License)
- PyTorch Image Models (timm) (Apache-2.0 License)
- CAVIS (MIT License)
- Mask2Former (Apache-2.0 License)
- Detectron2 (Apache-2.0 License)

