2 papers
cs.CV2025
Efficient High-Resolution Visual Representation Learning with State Space Model for Human Pose Estimation
Hao Zhang, Yongqiang Ma, Wenqi Shao +3
Capturing long-range dependencies while preserving high-resolution visual representations is crucial for dense prediction tasks such as human pose estimation. Vision Transformers (…
cs.CV2024
B-AVIBench: Towards Evaluating the Robustness of Large Vision-Language Model on Black-box Adversarial Visual-Instructions
Hao Zhang, Wenqi Shao, Hong Liu +5
Large Vision-Language Models (LVLMs) have shown significant progress in responding well to visual-instructions from users. However, these instructions, encompassing images and text…