ECCVW 2026 路 [Arxiv] [馃搫 Paper]
Brun贸 B. Englert, Gijs Dubbelman
Eindhoven University of Technology
Workshop: LIMIT: Representation Learning with Very Limited Resources
We present a controlled study of image and video self-supervised learning (SSL) for visual foundation models under limited resources. Using matched K700 data, a shared ViT-based architecture, input settings, and an 80k-update pretraining budget, we compare contrastive, reconstruction, feature-prediction, and diffusion objectives, including jointly trained image--video formulations.
We evaluate the resulting frozen representations on image classification, semantic segmentation, monocular depth estimation, action recognition, tracking, and relative camera-pose estimation. DINOv2-style pretraining provides the strongest overall performance in this resource-constrained setting, while combining DINOv2 with VideoMAE improves semantic image and video understanding. These gains come with lower tracking and camera-pose performance, exposing a tradeoff between semantic and geometric representation learning and motivating more balanced resource-efficient SSL objectives.
Available shortly
@inproceedings{englert26ssllimit,
author={{Englert, Brun\'o B.} and {Dubbelman, Gijs}},
title={A Controlled Study of Self-Supervised Image and Video Pretraining under Limited Resources},
booktitle={Computer Vision -- ECCV 2026 Workshops},
year={2026},
}