15 citations · 18 across the 3 of their papers we have counts for
3 papers
cs.CL2024
ZALM3: Zero-Shot Enhancement of Vision-Language Alignment via In-Context Information in Multi-Turn Multimodal Medical Dialogue
Zhangpu Li, Changhong Zou, Suxue Ma +13
The rocketing prosperity of large language models (LLMs) in recent years has boosted the prevalence of vision-language models (VLMs) in the medical sector. In our online medical co…
cs.CV2023★ 15 cited
HSTFormer: Hierarchical Spatial-Temporal Transformers for 3D Human Pose Estimation
Xiaoye Qian, Youbao Tang, Ning Zhang +4
Transformer-based approaches have been successfully proposed for 3D human pose estimation (HPE) from 2D pose sequence and achieved state-of-the-art (SOTA) performance. However, cur…
cs.CV2016★ 3 cited
AutoScaler: Scale-Attention Networks for Visual Correspondence
Shenlong Wang, Linjie Luo, Ning Zhang +1
Finding visual correspondence between local features is key to many computer vision problems. While defining features with larger contextual scales usually implies greater discrimi…