3 papers
cs.CV2026
AVTok: 1D Unified Tokenization for Holistic Audio-Video Generation
Kien T. Pham, I Chieh Chen, Qifeng Chen +1
Audio-video generation has recently gained unprecedented research attention, aiming to synthesize high-quality sounding video content with fine-grained synchronization and semantic…
cs.IR2025
ENCODE: Breaking the Trade-Off Between Performance and Efficiency in Long-Term User Behavior Modeling
Wenji Zhou, Yuhang Zheng, Yinfu Feng +5
Long-term user behavior sequences are a goldmine for businesses to explore users' interests to improve Click-Through Rate. However, it is very challenging to accurately capture use…
cs.GR2025
SpA2V: Harnessing Spatial Auditory Cues for Audio-driven Spatially-aware Video Generation
Kien T. Pham, Yingqing He, Yazhou Xing +2
Audio-driven video generation aims to synthesize realistic videos that align with input audio recordings, akin to the human ability to visualize scenes from auditory input. However…