activity
20222024
most citedVideo Activity Localisation with Uncertainties in Temporal Boundary

2 citations · 3 across the 7 of their papers we have counts for

collaborators

7 papers

cs.CV2024

MLLM as Video Narrator: Mitigating Modality Imbalance in Video Moment Retrieval

Weitong Cai, Jiabo Huang, Shaogang Gong +2

Video Moment Retrieval (VMR) aims to localize a specific temporal segment within an untrimmed long video given a natural language query. Existing methods often suffer from inadequa…

cs.CV2024

Hybrid-Learning Video Moment Retrieval across Multi-Domain Labels

Weitong Cai, Jiabo Huang, Shaogang Gong

Video moment retrieval (VMR) is to search for a visual temporal moment in an untrimmed raw video by a given text query description (sentence). Existing studies either start from co…

cs.SE2024

UniTSyn: A Large-Scale Dataset Capable of Enhancing the Prowess of Large Language Models for Program Testing

Yifeng He, Jiabo Huang, Yuyang Rong +3

The remarkable capability of large language models (LLMs) in generating high-quality code has drawn increasing attention in the software testing community. However, existing code L…

cs.SE2023

Code Representation Pre-training with Complements from Program Executions

Jiabo Huang, Jianyu Zhao, Yuyang Rong +3

Large language models (LLMs) for natural language processing have been grafted onto programming language modeling for advancing code intelligence. Although it can be represented in…

cs.CV2023

Zero-Shot Video Moment Retrieval from Frozen Vision-Language Models

Dezhao Luo, Jiabo Huang, Shaogang Gong +2

Accurate video moment retrieval (VMR) requires universal visual-textual correlations that can handle unknown vocabulary and unseen scenes. However, the learned correlations are lik…

cs.CV20231 cited

Towards Generalisable Video Moment Retrieval: Visual-Dynamic Injection to Image-Text Pre-Training

Dezhao Luo, Jiabo Huang, Shaogang Gong +2

The correlation between the vision and text is essential for video moment retrieval (VMR), however, existing methods heavily rely on separate pre-training feature extractors for vi…