collaborators

10 papers

cs.CV2026

Benchmarking Single-Factor Physical Video-to-Audio Generation

Tingle Li, Siddharth Gururani, Kevin J. Shih +6

Generative video-to-audio (V2A) models produce highly plausible soundtracks, but it remains unclear whether they capture the underlying physical processes. Existing evaluations emp…

cs.CL2026

S-MARC: Causal Streaming Reasoning for Full-Duplex Conversational Behavior Modeling

Dingkun Zhou, Shuchang Pan, Jiachen Lian +11

Human conversation is organized by an implicit chain of thought and manifests as temporally structured conversational behaviors. Capturing this perceptual pathway is critical for b…

cs.MM2026

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues

Dingkun Zhou, Krish Patel, Ajay Kankipati +14

Emotions conveyed through voice and face shape engagement and context in human AI interaction. Despite rapid progress in omni modal large language models, the holistic evaluation o…

cs.CL2025

Enabling Conversational Behavior Reasoning Capabilities in Full-Duplex Speech

Shuchang Pan, Siddharth Banerjee, Dhruv Hebbar +9

Human conversation is organized by an implicit chain of thoughts that manifests as timed speech acts. Capturing this causal pathway is key to building natural full-duplex interacti…

cs.CV2025

Schrodinger Audio-Visual Editor: Object-Level Audiovisual Removal

Weihan Xu, Kan Jen Cheng, Koichi Saito +10

Joint editing of audio and visual content is crucial for precise and controllable content creation. This new task poses challenges due to the limitations of paired audio-visual dat…

cs.RO2025

The Sound of Simulation: Learning Multimodal Sim-to-Real Robot Policies with Generative Audio

Renhao Wang, Haoran Geng, Tingle Li +7

Robots must integrate multiple sensory modalities to act effectively in the real world. Yet, learning such multimodal policies at scale remains challenging. Simulation offers a via…