3 papers
cs.LG2026
FuseAlign: Forced Alignment in the Wild
Mithilesh Vaidya, Stephen Bailey, Sumukh Badam +2
Word-level forced alignment estimates when each transcript word occurs in an audio recording. It underpins text-based media editing, subtitling, speech-data curation, and phonetic…
cs.CV2026
Goodbye Drift: Anchored Tree Sampling for Long-Horizon Video-to-Video Generation
Matthew Bendel, Stephen W. Bailey, Mithilesh Vaidya +2
Long-horizon video generation suffers from two intertwined issues. First, there is drift, where video quality degrades over time. Second, there are continuity issues which manifest…
eess.AS2026
PoDAR: Power-Decoupled Audio Representation for Generative Modeling
Alejandro Luebs, Mithilesh Vaidya, Ishaan Kumar +5
The performance of audio latent diffusion models is primarily governed by generator expressivity and the modelability of the underlying latent space. While recent research has focu…