2 papers
cs.AI2026
SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models
Tianyu Xie, Jinfa Huang, Yuexiao Ma +11
Omni-modal large language models (OLMs) redefine human-machine interaction by natively integrating audio, vision, and text. However, existing OLM benchmarks remain anchored to stat…
cs.RO2025
Embodied Tree of Thoughts: Deliberate Manipulation Planning with Embodied World Model
Wenjiang Xu, Cindy Wang, Rui Fang +6
World models have emerged as a pivotal component in robot manipulation planning, enabling agents to predict future environmental states and reason about the consequences of actions…