4 papers
AVTrack: Audio-Visual Tracking in Human-centric Complex Scenes
Yaoting Wang, Yun Zhou, Zipei Zhang +1
Audio-visual speaker tracking aims to localize and track active speakers by leveraging auditory and visual cues, enabling fine-grained, human-centric scene understanding. This capa…
FMBench: Adaptive Large Language Model Output Formatting
Yaoting Wang, Yun Zhou, Henghui Ding
Producing outputs that satisfy both semantic intent and format constraints is essential for deploying large language models in user-facing and system-integrated workflows. In this…
Audit After Segmentation: Reference-Free Mask Quality Assessment for Language-Referred Audio-Visual Segmentation
Jinxing Zhou, Yanghao Zhou, Yaoting Wang +5
Language-referred audio-visual segmentation (Ref-AVS) aims to segment target objects described by natural language by jointly reasoning over video, audio, and text. Beyond generati…
Ref-SAM3D: Bridging SAM3D with Text for Reference 3D Reconstruction
Yun Zhou, Yaoting Wang, Guangquan Jie +2
SAM3D has garnered widespread attention for its strong 3D object reconstruction capabilities. However, a key limitation remains: SAM3D cannot reconstruct specific objects referred…