Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Agentic Active Omni-Modal Perception for Multi-Hop Audio-Visual Reasoning
Ke Xu, Yuhao Wang, Ziyang Cheng +3
Multi-hop audio-visual reasoning remains challenging for Omni-LLMs, as relevant evidence is often sparse, temporally dispersed, and distributed across both audio and visual streams…
cs.AI2026
SceneCode: Executable World Programs for Editable Indoor Scenes with Articulated Objects
Puyi Wang, Yuhao Wang, Linjie Li +4
Indoor scene synthesis underpins embodied AI, robotic manipulation, and simulation-based policy evaluation, where a useful scene must specify not only what the environment looks li…