7 papers
PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation
Tianxin Xie, Wentao Lei, Kai Jiang +27
Text-to-audio-video (T2AV) generation is central to applications such as filmmaking and world modeling. However, current models often fail to produce physically plausible sounds. P…
Cued-Agent: A Collaborative Multi-Agent System for Automatic Cued Speech Recognition
Guanjie Huang, Danny H. K. Tsang, Shan Yang +2
Cued Speech (CS) is a visual communication system that combines lip-reading with hand coding to facilitate communication for individuals with hearing impairments. Automatic CS Reco…
Underwater Camouflaged Object Tracking Meets Vision-Language SAM2
Chunhui Zhang, Li Liu, Guanjie Huang +5
Over the past decade, significant progress has been made in visual object tracking, largely due to the availability of large-scale datasets. However, these datasets have primarily…
Lend a Hand: Semi Training-Free Cued Speech Recognition via MLLM-Driven Hand Modeling for Barrier-free Communication
Guanjie Huang, Danny Hin Kwok Tsang, Li Liu
Cued Speech (CS) is an innovative visual communication system that integrates lip-reading with hand coding, designed to enhance effective communication for individuals with hearing…
Segment Anything for Videos: A Systematic Survey
Chunhui Zhang, Yawen Cui, Weilin Lin +4
The recent wave of foundation models has witnessed tremendous success in computer vision (CV) and beyond, with the segment anything model (SAM) having sparked a passion for explori…
A Comprehensive Survey on Human Video Generation: Challenges, Methods, and Insights
Wentao Lei, Jinting Wang, Fengji Ma +2
Human video generation is a dynamic and rapidly evolving task that aims to synthesize 2D human body video sequences with generative models given control conditions such as text, au…