137 citations · 252 across the 16 of their papers we have counts for
24 papers
LVLM-eHub: A Comprehensive Evaluation Benchmark for Large Vision-Language Models
Peng Xu, Wenqi Shao, Kaipeng Zhang +7
Large Vision-Language Models (LVLMs) have recently played a dominant role in multimodal vision-language learning. Despite the great success, it lacks a holistic evaluation of their…
Perceive, Ground, Reason, and Act: A Benchmark for General-purpose Visual Representation
Jiangyong Huang, William Yicheng Zhu, Baoxiong Jia +4
Current computer vision models, unlike the human visual system, cannot yet achieve general-purpose visual understanding. Existing efforts to create a general vision model are limit…
HUMANISE: Language-conditioned Human Motion Generation in 3D Scenes
Zan Wang, Yixin Chen, Tengyu Liu +3
Learning to generate diverse scene-aware and goal-oriented human motions in 3D scenes remains challenging due to the mediocre characteristics of the existing datasets on Human-Scen…
EgoTaskQA: Understanding Human Tasks in Egocentric Videos
Baoxiong Jia, Ting Lei, Song-Chun Zhu +1
Understanding human tasks through video observations is an essential capability of intelligent agents. The challenges of such capability lie in the difficulty of generating a detai…
Infrared Invisible Clothing:Hiding from Infrared Detectors at Multiple Angles in Real World
Xiaopei Zhu, Zhanhao Hu, Siyuan Huang +2
Thermal infrared imaging is widely used in body temperature measurement, security monitoring, and so on, but its safety research attracted attention only in recent years. We propos…
PartAfford: Part-level Affordance Discovery from 3D Objects
Chao Xu, Yixin Chen, He Wang +3
Understanding what objects could furnish for humans-namely, learning object affordance-is the crux to bridge perception and action. In the vision community, prior work primarily fo…