6 papers
GTR-Bench: Evaluating Geo-Temporal Reasoning in Vision-Language Models
Qinghongbing Xie, Zhaoyuan Xia, Feng Zhu +4
Recently spatial-temporal intelligence of Visual-Language Models (VLMs) has attracted much attention due to its importance for autonomous driving, embodied AI and general AI. Exist…
Re-Aligning Language to Visual Objects with an Agentic Workflow
Yuming Chen, Jiangyan Feng, Haodong Zhang +6
Language-based object detection (LOD) aims to align visual objects with language expressions. A large amount of paired data is utilized to improve LOD model generalizations. During…
Instruct-ReID++: Towards Universal Purpose Instruction-Guided Person Re-identification
Weizhen He, Yiheng Deng, Yunfeng Yan +7
Human intelligence can retrieve any person according to both visual and language descriptions. However, the current computer vision community studies specific person re-identificat…
VisionTraj: A Noise-Robust Trajectory Recovery Framework based on Large-scale Camera Network
Zhishuai Li, Ziyue Li, Xiaoru Hu +5
Trajectory recovery based on the snapshots from the city-wide multi-camera network facilitates urban mobility sensing and driveway optimization. The state-of-the-art solutions devo…
Hulk: A Universal Knowledge Translator for Human-Centric Tasks
Yizhou Wang, Yixuan Wu, Weizhen He +8
Human-centric perception tasks, e.g., pedestrian detection, skeleton-based action recognition, and pose estimation, have wide industrial applications, such as metaverse and sports…
InstructDET: Diversifying Referring Object Detection with Generalized Instructions
Ronghao Dang, Jiangyan Feng, Haodong Zhang +8
We propose InstructDET, a data-centric method for referring object detection (ROD) that localizes target objects based on user instructions. While deriving from referring expressio…