10 papers
Benchmarking Single-Factor Physical Video-to-Audio Generation
Tingle Li, Siddharth Gururani, Kevin J. Shih +6
Generative video-to-audio (V2A) models produce highly plausible soundtracks, but it remains unclear whether they capture the underlying physical processes. Existing evaluations emp…
S-MARC: Causal Streaming Reasoning for Full-Duplex Conversational Behavior Modeling
Dingkun Zhou, Shuchang Pan, Jiachen Lian +11
Human conversation is organized by an implicit chain of thought and manifests as temporally structured conversational behaviors. Capturing this perceptual pathway is critical for b…
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues
Dingkun Zhou, Krish Patel, Ajay Kankipati +14
Emotions conveyed through voice and face shape engagement and context in human AI interaction. Despite rapid progress in omni modal large language models, the holistic evaluation o…
Enabling Conversational Behavior Reasoning Capabilities in Full-Duplex Speech
Shuchang Pan, Siddharth Banerjee, Dhruv Hebbar +9
Human conversation is organized by an implicit chain of thoughts that manifests as timed speech acts. Capturing this causal pathway is key to building natural full-duplex interacti…
Schrodinger Audio-Visual Editor: Object-Level Audiovisual Removal
Weihan Xu, Kan Jen Cheng, Koichi Saito +10
Joint editing of audio and visual content is crucial for precise and controllable content creation. This new task poses challenges due to the limitations of paired audio-visual dat…
The Sound of Simulation: Learning Multimodal Sim-to-Real Robot Policies with Generative Audio
Renhao Wang, Haoran Geng, Tingle Li +7
Robots must integrate multiple sensory modalities to act effectively in the real world. Yet, learning such multimodal policies at scale remains challenging. Simulation offers a via…