2 citations · 2 across the 2 of their papers we have counts for
2 papers
cs.CV2026
EgoMonth: A Month-Level Egocentric Video Benchmark for Long-Term Spatiotemporal Memory
Weitao Chen, Hu Jiaxin, Xie Tianyidan +15
Recent advances in Multimodal Large Language Models (MLLMs) have led to substantial progress in video understanding, accompanied by a growing number of long video benchmarks. Howev…
cs.CV2023★ 2 cited
Sam-Guided Enhanced Fine-Grained Encoding with Mixed Semantic Learning for Medical Image Captioning
Zhenyu Zhang, Benlu Wang, Weijie Liang +5
With the development of multimodality and large language models, the deep learning-based technique for medical image captioning holds the potential to offer valuable diagnostic rec…