2 citations · 2 across the 3 of their papers we have counts for
3 papers
cs.CV2025
EvdCLIP: Improving Vision-Language Retrieval with Entity Visual Descriptions from Large Language Models
GuangHao Meng, Sunan He, Jinpeng Wang +7
Vision-language retrieval (VLR) has attracted significant attention in both academia and industry, which involves using text (or images) as queries to retrieve corresponding images…
cs.CV2024★ 2 cited
Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning
Bingchen Zhao, Yongshuo Zong, Letian Zhang +1
The advancement of large language models (LLMs) has significantly broadened the scope of applications in natural language processing, with multi-modal LLMs extending these capabili…
cs.CV2023
EPose: Energy-Efficient Edge-assisted Multi-camera System for Multi-human 3D Pose Estimation
Letian Zhang, Jie Xu
Multi-human 3D pose estimation plays a key role in establishing a seamless connection between the real world and the virtual world. Recent efforts adopted a two-stage framework tha…