Skip to content

About

MVFGA preprocessing pipeline

Resources

Stars

6 stars

Watchers

0 watching

Forks

Repository files navigation

MVFGA Dataset Preprocessing Pipeline

An end-to-end preprocessing toolkit for multi-view face and gesture animation datasets.

Project Page Paper Dataset License: MIT


Overview

This repository provides the preprocessing pipeline used for MVFGA (Multi-View Face and Gesture Animation with Dynamic Gaussians). It contains scripts for preparing synchronized multi-view recordings, including video trimming, camera calibration, background matting, body segmentation, and landmark detection.

Pipeline

MVFGA Preprocessing Pipeline

The preprocessing workflow consists of the following stages:

Stage Task Description
1 Video synchronization and trimming Aligns multi-camera recordings and trims them to a common temporal range.
2 Camera calibration Estimates intrinsic and extrinsic camera parameters from checkerboard recordings.
3 Background matting Extracts foreground subjects using BiRefNet.
4 Body segmentation Produces human-body segmentation masks using Sapiens.
5 Landmark detection Extracts MediaPipe whole-body 2D landmarks through the EasyMoCap wrapper.

Installation

Create and activate the Conda environment provided with the repository:

chmod +x scripts/install_conda.sh
./scripts/install_conda.sh

sudo apt-get update
sudo apt-get install -y git

conda env create -f environment.yml
conda activate mvfga
export CUDA_HOME="$CONDA_PREFIX"

Install PyTorch3D:

pip install "git+https://github.com/facebookresearch/pytorch3d.git@stable"

Install HaMeR and ViTPose:

cd hamer
pip install -e ".[all]"
pip install -v -e third-party/ViTPose
cd ..

Install EasyMoCap in development mode:

cd Easymocap
python setup.py develop
cd ..

Review environment.yml for the complete list of dependencies and version requirements.

Usage

Synchronize multi-view videos

The synchronization interface lets you align all camera streams to the same start or end time. You can move through frames for all videos simultaneously or adjust an individual stream through the GUI controls.

cd sync_streams
python3 multiview_sync.py --video_dir "D:/Datacapture/Extrinsics_1/videos"

Run the preprocessing pipeline

python main.py \
  --root_dir /home/vippin/thesis/extra \
  --output /home/vippin/thesis/extra/demo_mvfga \
  --sapiens \
  --calibrate \
  --background_matting \
  --annots

The example above enables camera calibration, background matting, Sapiens body segmentation, and landmark annotation generation.

Processing Stages

1. Trim videos

The trimming stage processes synchronized video streams and crops each recording to its selected start and end frames.

2. Camera calibration

The calibration stage estimates intrinsic and extrinsic camera parameters using checkerboard recordings.

For additional details, see the EasyMoCap calibration documentation.

3. Background matting

The background-matting stage separates the subject from the background using BiRefNet.

For additional details, see the BiRefNet documentation.

4. Body segmentation

The body-segmentation stage generates human segmentation masks using Sapiens.

For additional details, see the Sapiens segmentation documentation.

5. MediaPipe landmark detection

EasyMoCap provides a wrapper for extracting MediaPipe whole-body 2D keypoints from the processed sequences.

Logging

Each pipeline run writes detailed progress information, warnings, and errors to:

output.log

Citation

Please cite our paper when using MVFGA or this preprocessing pipeline in your research:

@article{javanmardi2026multi,
  title = {Multi-View Face and Gesture Animation with Dynamic Gaussians},
  author = {Javanmardi, Alireza and Jeetmal, Vippin Kumar and Millerdurai, Christen and Pagani, Alain and Stricker, Didier},
  journal = {arXiv preprint arXiv:2608.04722},
  year = {2026}
}

Referenced methods and tools

Sapiens
@article{khirodkar2024sapiens,
  title   = {Sapiens: Foundation for Human Vision Models},
  author  = {Khirodkar, Rawal and Bagautdinov, Timur and Martinez, Julieta and Zhaoen, Su and James, Austin and Selednik, Peter and Anderson, Stuart and Saito, Shunsuke},
  journal = {arXiv preprint arXiv:2408.12569},
  year    = {2024}
}
Background Matting — BiRefNet
@article{zheng2024birefnet,
  title   = {Bilateral Reference for High-Resolution Dichotomous Image Segmentation},
  author  = {Zheng, Peng and Gao, Dehong and Fan, Deng-Ping and Liu, Li and Laaksonen, Jorma and Ouyang, Wanli and Sebe, Nicu},
  journal = {CAAI Artificial Intelligence Research},
  volume  = {3},
  pages   = {9150038},
  year    = {2024}
}
EasyMoCap
@misc{easymocap,
  title        = {EasyMoCap - Make human motion capture easier.},
  howpublished = {GitHub},
  year         = {2021},
  url          = {https://github.com/zju3dv/EasyMocap}
}
Metrical Face Tracker
@proceedings{MICA:ECCV2022,
  author  = {Zielonka, Wojciech and Bolkart, Timo and Thies, Justus},
  title   = {Towards Metrical Reconstruction of Human Faces},
  journal = {European Conference on Computer Vision},
  year    = {2022}
}
HaMeR
@inproceedings{pavlakos2024reconstructing,
  title     = {Reconstructing Hands in 3{D} with Transformers},
  author    = {Pavlakos, Georgios and Shan, Dandan and Radosavovic, Ilija and Kanazawa, Angjoo and Fouhey, David and Malik, Jitendra},
  booktitle = {CVPR},
  year      = {2024}
}

Acknowledgements

This work was partially funded by the Horizon Europe programme under the IRIS-XR project, Grant Agreement No. 101298672.

Contributing

Contributions are welcome. Please fork the repository, create a focused branch, and submit a pull request with a clear description of your changes.

License

This project is licensed under the MIT License.

About

MVFGA preprocessing pipeline

Resources

Stars

6 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages