4 papers
SPARTUN3D: Situated Spatial Understanding of 3D World in Large Language Models
Yue Zhang, Zhiyang Xu, Ying Shen +2
Integrating the 3D world into large language models (3D-based LLMs) has been a promising research direction for 3D scene understanding. However, current 3D-based LLMs fall short in…
Vision-and-Language Navigation Today and Tomorrow: A Survey in the Era of Foundation Models
Yue Zhang, Ziqiao Ma, Jialu Li +6
Vision-and-Language Navigation (VLN) has gained increasing attention over recent years and many approaches have emerged to advance their development. The remarkable achievements of…
Narrowing the Gap between Vision and Action in Navigation
Yue Zhang, Parisa Kordjamshidi
The existing methods for Vision and Language Navigation in the Continuous Environment (VLN-CE) commonly incorporate a waypoint predictor to discretize the environment. This simplif…
Common Sense Reasoning for Deepfake Detection
Yue Zhang, Ben Colman, Xiao Guo +2
State-of-the-art deepfake detection approaches rely on image-based features extracted via neural networks. While these approaches trained in a supervised manner extract likely fake…