1 citations · 1 across the 13 of their papers we have counts for
9 papers · 1 filter
Self-Paced Learning for Images of Antinuclear Antibodies
Yiyang Jiang, Guangwu Qian, Jiaxin Wu +4
Antinuclear antibody (ANA) testing is a crucial method for diagnosing autoimmune disorders, including lupus, Sjögren's syndrome, and scleroderma. Despite its importance, manual ANA…
Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning
Dayong Liang, Changmeng Zheng, Zhiyuan Wen +3
Traditional scene graphs primarily focus on spatial relationships, limiting vision-language models' (VLMs) ability to reason about complex interactions in visual scenes. This paper…
Mean of Means: Human Localization with Calibration-free and Unconstrained Camera Settings (extended version)
Tianyi Zhang, Wengyu Zhang, Xulu Zhang +4
Accurate human localization is crucial for various applications, especially in the Metaverse era. Existing high precision solutions rely on expensive, tag-dependent hardware, while…
PolySmart @ TRECVid 2024 Video Captioning (VTT)
Jiaxin Wu, Wengyu Zhang, Xiao-Yong Wei +1
In this paper, we present our methods and results for the Video-To-Text (VTT) task at TRECVid 2024, exploring the capabilities of Vision-Language Models (VLMs) like LLaVA and LLaVA…
PolySmart @ TRECVid 2024 Medical Video Question Answering
Jiaxin Wu, Yiyang Jiang, Xiao-Yong Wei +1
Video Corpus Visual Answer Localization (VCVAL) includes question-related video retrieval and visual answer localization in the videos. Specifically, we use text-to-text retrieval…
Mean of Means: A 10-dollar Solution for Human Localization with Calibration-free and Unconstrained Camera Settings
Tianyi Zhang, Wengyu Zhang, Xulu Zhang +4
Accurate human localization is crucial for various applications, especially in the Metaverse era. Existing high precision solutions rely on expensive, tag-dependent hardware, while…