From the 2 of 9 linked papers with an AI index.
9 papers
RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination
Haotian Liang, Mingkang Chen, Yufei Huang +27
The paper introduces RxBrain, a foundation model that jointly reasons over language and visual inputs to create embodied plans, using a multimodal Mixture-of-Transformers architect…
ProAct: A Benchmark and Multimodal Framework for Structure-Aware Proactive Response
Xiaomeng Zhu, Fengming Zhu, Weijie Zhou +8
The paper introduces ProAct-75, a benchmark of 75 proactive tasks with step‑level annotations and task graphs, and presents ProAct-Helper, a multimodal LLM that uses these graphs f…
UniBYD: A Unified Framework for Learning Robotic Manipulation Across Embodiments Beyond Imitation of Human Demonstrations
Tingyu Yuan, Biaoliang Guan, Wen Ye +8
In embodied intelligence, the embodiment gap between robotic and human hands brings significant challenges for learning from human demonstrations. Although some studies have attemp…
Listening with the Eyes: Benchmarking Egocentric Co-Speech Grounding across Space and Time
Weijie Zhou, Xuantang Xiong, Zhenlin Hu +6
In situated collaboration, speakers often use intentionally underspecified deictic commands (e.g., ``pass me \textit{that}''), whose referent becomes identifiable only by aligning…
SteerEval: A Framework for Evaluating Steerability with Natural Language Profiles for Recommendation
Joyce Zhou, Weijie Zhou, Doug Turnbull +1
Natural-language user profiles have recently attracted attention not only for improved interpretability, but also for their potential to make recommender systems more steerable. By…
PhysVLM-AVR: Active Visual Reasoning for Multimodal Large Language Models in Physical Environments
Weijie Zhou, Xuantang Xiong, Yi Peng +5
Visual reasoning in multimodal large language models (MLLMs) has primarily been studied in static, fully observable settings, limiting their effectiveness in real-world environment…