3 papers
cs.CV2026
Bridging the Visual-to-Physical Gap: Physically Aligned Representations for Fall Risk Analysis
Xianqi Zhang
Vision-based fall analysis has advanced rapidly, but a key bottleneck remains: visually similarmotions can correspond to very different physical outcomes because small differences…
cs.CV2025
Region-Level Context-Aware Multimodal Understanding
Hongliang Wei, Xianqi Zhang, Xingtao Wang +2
Despite significant progress, existing research on Multimodal Large Language Models (MLLMs) mainly focuses on general visual understanding, overlooking the ability to integrate tex…
cs.RO2025
FLAM: Foundation Model-Based Body Stabilization for Humanoid Locomotion and Manipulation
Xianqi Zhang, Hongliang Wei, Wenrui Wang +3
Humanoid robots have attracted significant attention in recent years. Reinforcement Learning (RL) is one of the main ways to control the whole body of humanoid robots. RL enables a…