40 papers
Automated Discovery Has No Universally Superior Harness
Akshat Gupta, Jermaine Lei, Alexander Lu +2
Autonomous discovery systems such as OpenEvolve and TTT-Discover are often used as general-purpose harnesses. However, in practice these are composite systems combining several des…
Speech World Model: Causal State-Action Planning with Explicit Reasoning for Speech
Xuanru Zhou, Jiachen Lian, Henry Hong +2
Current speech-language models (SLMs) typically use a cascade of speech encoder and large language model, treating speech understanding as a single black box. They analyze the cont…
Benchmarking Single-Factor Physical Video-to-Audio Generation
Tingle Li, Siddharth Gururani, Kevin J. Shih +6
Generative video-to-audio (V2A) models produce highly plausible soundtracks, but it remains unclear whether they capture the underlying physical processes. Existing evaluations emp…
S-MARC: Causal Streaming Reasoning for Full-Duplex Conversational Behavior Modeling
Dingkun Zhou, Shuchang Pan, Jiachen Lian +11
Human conversation is organized by an implicit chain of thought and manifests as temporally structured conversational behaviors. Capturing this perceptual pathway is critical for b…
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues
Dingkun Zhou, Krish Patel, Ajay Kankipati +14
Emotions conveyed through voice and face shape engagement and context in human AI interaction. Despite rapid progress in omni modal large language models, the holistic evaluation o…
HASS: Hierarchical Simulation of Logopenic Aphasic Speech for Scalable PPA Detection
Harrison Li, Kevin Wang, Cheol Jun Cho +13
Building a diagnosis model for primary progressive aphasia (PPA) has been challenging due to the data scarcity. Collecting clinical data at scale is limited by the high vulnerabili…