activity
20212025
most citedRNet:Relation-embedded Representation Reconstruction Network for Change Captioning

1 citations · 1 across the 3 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2025

Collaborative Temporal Consistency Learning for Point-supervised Natural Language Video Localization

Zhuo Tao, Liang Li, Qi Chen +5

Natural language video localization (NLVL) is a crucial task in video understanding that aims to localize the target moment in videos specified by a given language description. Rec…

cs.CV2024

Query-centric Audio-Visual Cognition Network for Moment Retrieval, Segmentation and Step-Captioning

Yunbin Tu, Liang Li, Li Su +1

Video has emerged as a favored multimedia format on the internet. To better gain video contents, a new topic HIREST is presented, including video retrieval, moment retrieval, momen…

cs.CV2024

Distractors-Immune Representation Learning with Cross-modal Contrastive Regularization for Change Captioning

Yunbin Tu, Liang Li, Li Su +2

Change captioning aims to succinctly describe the semantic change between a pair of similar images, while being immune to distractors (illumination and viewpoint changes). Under th…

cs.CV2024

Context-aware Difference Distilling for Multi-change Captioning

Yunbin Tu, Liang Li, Li Su +3

Multi-change captioning aims to describe complex and coupled changes within an image pair in natural language. Compared with single-change captioning, this task requires the model…

cs.CV2023

Self-supervised Cross-view Representation Reconstruction for Change Captioning

Yunbin Tu, Liang Li, Li Su +3

Change captioning aims to describe the difference between a pair of similar images. Its key challenge is how to learn a stable difference representation under pseudo changes caused…