49 citations · 72 across the 8 of their papers we have counts for
10 papers · 1 filter
SE-VLN: A Self-Evolving Vision-Language Navigation Framework Based on Multimodal Large Language Models
Xiangyu Dong, Haoran Zhao, Jiang Gao +5
Recent advances in vision-language navigation (VLN) were mainly attributed to emerging large language models (LLMs). These methods exhibited excellent generalization capabilities i…
Cross-modal Semantic Enhanced Interaction for Image-Sentence Retrieval
Xuri Ge, Fuhai Chen, Songpei Xu +2
Image-sentence retrieval has attracted extensive research attention in multimedia and computer vision due to its promising application. The key issue lies in jointly learning the v…
Factored Attention and Embedding for Unstructured-view Topic-related Ultrasound Report Generation
Fuhai Chen, Rongrong Ji, Chengpeng Dai +4
Echocardiography is widely used to clinical practice for diagnosis and treatment, e.g., on the common congenital heart defects. The traditional manual manipulation is error-prone d…
Global2Local: A Joint-Hierarchical Attention for Video Captioning
Chengpeng Dai, Fuhai Chen, Xiaoshuai Sun +3
Recently, automatic video captioning has attracted increasing attention, where the core challenge lies in capturing the key semantic items, like objects and actions as well as thei…
Differentiated Relevances Embedding for Group-based Referring Expression Comprehension
Fuhai Chen, Xuri Ge, Xiaoshuai Sun +4
The key of referring expression comprehension lies in capturing the cross-modal visual-linguistic relevance. Existing works typically model the cross-modal relevance in each image,…
Weakly-Supervised Dense Action Anticipation
Haotong Zhang, Fuhai Chen, Angela Yao
Dense anticipation aims to forecast future actions and their durations for long horizons. Existing approaches rely on fully-labelled data, i.e. sequences labelled with all future a…