4 papers
Agentic Active Omni-Modal Perception for Multi-Hop Audio-Visual Reasoning
Ke Xu, Yuhao Wang, Ziyang Cheng +3
Multi-hop audio-visual reasoning remains challenging for Omni-LLMs, as relevant evidence is often sparse, temporally dispersed, and distributed across both audio and visual streams…
Evolving Robustness--Exploration Trade-off in Online Reinforcement Learning via Quantile Bayesian Risk MDPs
Meichen Song, Yuhao Wang, Enlu Zhou
In online reinforcement learning, data scarcity creates epistemic uncertainty that makes robustness important early in learning, whereas sufficient exploration is needed to learn t…
StepAudio 2.5 Technical Report
Bin Lin, Bo Zhao, Boyong Wu +98
Unified audio-language modeling has emerged as a prominent trend in modern speech systems, promising to bring the reasoning capabilities of large language models to auditory tasks.…
SceneCode: Executable World Programs for Editable Indoor Scenes with Articulated Objects
Puyi Wang, Yuhao Wang, Linjie Li +4
Indoor scene synthesis underpins embodied AI, robotic manipulation, and simulation-based policy evaluation, where a useful scene must specify not only what the environment looks li…