2 papers
cs.CL2025
VIVA+: Human-Centered Situational Decision-Making
Zhe Hu, Yixiao Ren, Guanzhong Liu +2
Multimodal Large Language Models (MLLMs) show promising results for embodied agents in operating meaningfully in complex, human-centered environments. Yet, evaluating their capacit…
cs.CL2025
Praxis-VLM: Vision-Grounded Decision Making via Text-Driven Reinforcement Learning
Zhe Hu, Jing Li, Zhongzhu Pu +2
Vision Language Models exhibit impressive performance for various tasks, yet they often lack the sophisticated situational reasoning required for complex decision-making. This pape…