2 papers
cs.CV2025
LongCaptioning: Unlocking the Power of Long Video Caption Generation in Large Multimodal Models
Hongchen Wei, Zhihong Tan, Yaosi Hu +2
Large Multimodal Models (LMMs) have demonstrated exceptional performance in video captioning tasks, particularly for short videos. However, as the length of the video increases, ge…
eess.IV2025
Remote Sensing Semantic Segmentation Quality Assessment based on Vision Language Model
Huiying Shi, Zhihong Tan, Zhihan Zhang +4
The complexity of scenes and variations in image quality result in significant variability in the performance of semantic segmentation methods of remote sensing imagery (RSI) in su…