2 papers
eess.AS2026
PoDAR: Power-Disentangled Audio Representation for Generative Modeling
Alejandro Luebs, Mithilesh Vaidya, Ishaan Kumar +5
The performance of audio latent diffusion models is primarily governed by generator expressivity and the modelability of the underlying latent space. While recent research has focu…
cs.CV2026
UpstreamQA: A Modular Framework for Explicit Reasoning on Video Question Answering Tasks
Jason Nguyen, Ameet Rao, Alexander Chang +2
Video Question Answering (VideoQA) demands models that jointly reason over spatial, temporal, and linguistic cues. However, the task's inherent complexity often requires multi-step…