97 citations · 239 across the 14 of their papers we have counts for
6 papers · 1 filter
LVLM-eHub: A Comprehensive Evaluation Benchmark for Large Vision-Language Models
Peng Xu, Wenqi Shao, Kaipeng Zhang +7
Large Vision-Language Models (LVLMs) have recently played a dominant role in multimodal vision-language learning. Despite the great success, it lacks a holistic evaluation of their…
Adaptive Multi-Modal Cross-Entropy Loss for Stereo Matching
Peng Xu, Zhiyu Xiang, Chenyu Qiao +2
Despite the great success of deep learning in stereo matching, recovering accurate disparity maps is still challenging. Currently, L1 and cross-entropy are the two most widely used…
SAM Struggles in Concealed Scenes -- Empirical Study on Segment Anything
Ge-Peng Ji, Deng-Ping Fan, Peng Xu +3
Segmenting anything is a ground-breaking step toward artificial general intelligence, and the Segment Anything Model (SAM) greatly fosters the foundation models for computer vision…
How to Understand Masked Autoencoders
Shuhao Cao, Peng Xu, David A. Clifton
"Masked Autoencoders (MAE) Are Scalable Vision Learners" revolutionizes the self-supervised learning method in that it not only achieves the state-of-the-art for image pre-training…
SPIN: A High Speed, High Resolution Vision Dataset for Tracking and Action Recognition in Ping Pong
Steven Schwarcz, Peng Xu, David D'Ambrosio +4
We introduce a new high resolution, high frame rate stereo video dataset, which we call SPIN, for tracking and action recognition in the game of ping pong. The corpus consists of p…
Boosting Image Recognition with Non-differentiable Constraints
Xuan Li, Yuchen Lu, Peng Xu +3
In this paper, we study the problem of image recognition with non-differentiable constraints. A lot of real-life recognition applications require a rich output structure with deter…