Scripts for "Montreal Forced Aligner and the state of speech-to-text alignment in 2026" to be presented at Interspeech 2026
Preprint of paper available on here
- MFA: 3.4
- Benchmarks were done with in development version of 3.4 in January 2026
- MAUS: WebMAUS, accessed January 2026
- MAPS: v0.5.0
- Not reported in the paper is the use of variants implemented in 0.5.0, which had similar performance on TIMIT, but much worse performance on Buckeye, so the original numbers were kept as the benchmark
- Korean Forced Aligner: Web portal, accessed January 2026
- WhisperX: 3.4.2
- Wav2Vec2: torchaudio 2.8.0
- Nemo Forced Aligner: 12251c3
- Julius: 1604011
- Adapted from pyJuliusAlign
- Bournemouth Forced Aligner: 0.1.7
Incomplete, but historical notes
- Korean forced aligner
- Upload and processing of Seoul Corpus files:
- Total time: ~6 hours
- initial s01 time: 11:43
- Final s20 time: 15:04
- initial s21 time: 16:15
- Final s22 time: 16:39
- initial s23 time: 10:27
- Final s40 time: 13:33
- Upload and processing of Seoul Corpus files: