1 citations · 1 across the 4 of their papers we have counts for
4 papers
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding
Yun Li, Zhe Liu, Yajing Kong +6
Applying Multimodal Large Language Models (MLLMs) to video understanding presents significant challenges due to the need to model temporal relations across frames. Existing approac…
VisionLLM-based Multimodal Fusion Network for Glottic Carcinoma Early Detection
Zhaohui Jin, Yi Shuai, Yongcheng Li +4
The early detection of glottic carcinoma is critical for improving patient outcomes, as it enables timely intervention, preserves vocal function, and significantly reduces the risk…
SAM-FNet: SAM-Guided Fusion Network for Laryngo-Pharyngeal Tumor Detection
Jia Wei, Yun Li, Meiyu Qiu +3
Laryngo-pharyngeal cancer (LPC) is a highly fatal malignant disease affecting the head and neck region. Previous studies on endoscopic tumor detection, particularly those leveragin…
Synthetic Hard Negative Samples for Contrastive Learning
Hengkui Dong, Xianzhong Long, Yun Li +1
Contrastive learning has emerged as an essential approach for self-supervised learning in visual representation learning. The central objective of contrastive learning is to maximi…