4 papers
MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation
Haitian Li, Yanghao Zhou, Heyan Huang +15
In recent years, Multi-Talker Audio-Video Generation (MTAVG) models have shown promising performance on fundamental metrics such as lip-sync and audio-visual alignment. However, th…
Zero-Effort Image-to-Music Generation: An Interpretable RAG-based VLM Approach
Zijian Zhao, Dian Jin, Zijing Zhou
Recently, Image-to-Music (I2M) generation has garnered significant attention, with potential applications in fields such as gaming, advertising, and multi-modal art creation. Howev…
Automatic Stage Lighting Control: Is it a Rule-Driven Process or Generative Task?
Zijian Zhao, Dian Jin, Zijing Zhou +1
Stage lighting is a vital component in live music performances, shaping an engaging experience for both musicians and audiences. In recent years, Automatic Stage Lighting Control (…
SimToken: A Simple Baseline for Referring Audio-Visual Segmentation
Dian Jin, Yanghao Zhou, Jinxing Zhou +3
Referring Audio-Visual Segmentation (Ref-AVS) aims to segment specific objects in videos based on natural language expressions involving audio, vision, and text information. This t…