3 papers
cs.CV2025
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning
Tianheng Qiu, Jingchun Gao, Jingyu Li +6
Intent-oriented controlled video captioning aims to generate targeted descriptions for specific targets in a video based on customized user intent. Current Large Visual Language Mo…
cs.CV2024
CLDA-YOLO: Visual Contrastive Learning Based Domain Adaptive YOLO Detector
Tianheng Qiu, Ka Lung Law, Guanghua Pan +4
Unsupervised domain adaptive (UDA) algorithms can markedly enhance the performance of object detectors under conditions of domain shifts, thereby reducing the necessity for extensi…
cs.CV2024
DSNet: A Novel Way to Use Atrous Convolutions in Semantic Segmentation
Zilu Guo, Liuyang Bian, Xuan Huang +3
Atrous convolutions are employed as a method to increase the receptive field in semantic segmentation tasks. However, in previous works of semantic segmentation, it was rarely empl…