From the 2 of 5 linked papers with an AI index.
5 papers
RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination
Haotian Liang, Mingkang Chen, Yufei Huang +27
The paper introduces RxBrain, a foundation model that jointly reasons over language and visual inputs to create embodied plans, using a multimodal Mixture-of-Transformers architect…
ProAct: A Benchmark and Multimodal Framework for Structure-Aware Proactive Response
Xiaomeng Zhu, Fengming Zhu, Weijie Zhou +8
The paper introduces ProAct-75, a benchmark of 75 proactive tasks with step‑level annotations and task graphs, and presents ProAct-Helper, a multimodal LLM that uses these graphs f…
Plan Right, Then Plan Tight: Symbolic RL for Efficient Embodied Reasoning
Xiangli Shi, Xiaomeng Zhu, Ye Tian +5
Embodied task planning asks an agent to turn a natural-language instruction into an executable sequence of actions in a physical scene, and is a building block for household, assis…
Language-Driven Object-Oriented Two-Stage Method for Scene Graph Anticipation
Xiaomeng Zhu, Changwei Wang, Haozhe Wang +2
A scene graph is a structured representation of objects and their spatio-temporal relationships in dynamic scenes. Scene Graph Anticipation (SGA) involves predicting future scene g…
A Computable Game-Theoretic Framework for Multi-Agent Theory of Mind
Fengming Zhu, Yuxin Pan, Xiaomeng Zhu +1
Originating in psychology, (ToM) has attracted significant attention across multiple research communities, especially logic, economics, and robotics. Most…