From the 3 of 9 linked papers with an AI index.
9 papers
RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination
Haotian Liang, Mingkang Chen, Yufei Huang +27
The paper introduces RxBrain, a foundation model that jointly reasons over language and visual inputs to create embodied plans, using a multimodal Mixture-of-Transformers architect…
Hy-Embodied-VLM-1.0: Efficient Physical-World Agents
Ziyi Wang, Xumin Yu, Yongming Rao +19
Building capable embodied agents requires not only multimodal perception and understanding, but also agentic capabilities for reasoning about actions, adapting to evolving situatio…
ProAct: A Benchmark and Multimodal Framework for Structure-Aware Proactive Response
Xiaomeng Zhu, Fengming Zhu, Weijie Zhou +8
The paper introduces ProAct-75, a benchmark of 75 proactive tasks with step‑level annotations and task graphs, and presents ProAct-Helper, a multimodal LLM that uses these graphs f…
Plan Right, Then Plan Tight: Symbolic RL for Efficient Embodied Reasoning
Xiangli Shi, Xiaomeng Zhu, Ye Tian +5
Embodied task planning asks an agent to turn a natural-language instruction into an executable sequence of actions in a physical scene, and is a building block for household, assis…
Improving Diffusion Generalization with Weak-to-Strong Segmented Guidance
Liangyu Yuan, Yufei Huang, Mingkun Lei +5
Diffusion models generate synthetic images through an iterative refinement process. However, the misalignment between the simulation-free objective and the iterative process often…
UltraLogic: Enhancing LLM Reasoning through Large-Scale Data Synthesis and Bipolar Float Reward
Yile Liu, Yixian Liu, Zongwei Li +7
While Large Language Models (LLMs) have demonstrated significant potential in natural language processing , complex general-purpose reasoning requiring multi-step logic, planning,…