1 paper · 1 filter
Xiaoda Yang, Yuxiang Liu, Shenzhou Gao +8
Modern vision-language models achieve strong performance in static perception, but remain limited in the complex spatiotemporal reasoning required for embodied, egocentric tasks. A…