2 papers
cs.CV2026
MAIN-VLA: Modeling Abstraction of Intention and eNvironment for Vision-Language-Action Models
Zheyuan Zhou, Liang Du, Zixun Sun +5
Despite significant progress in Visual-Language-Action (VLA), in highly complex and dynamic environments that involve real-time unpredictable interactions (such as 3D open worlds a…
cs.CV2025
Instruct-ReID: A Multi-purpose Person Re-identification Task with Instructions
Weizhen He, Yiheng Deng, Shixiang Tang +9
Human intelligence can retrieve any person according to both visual and language descriptions. However, the current computer vision community studies specific person re-identificat…