collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2026

Not All Inputs Are Valid: Towards Open-Set Video Moment Retrieval Using Language

Xiang Fang, Wanlong Fang, Daizong Liu +8

Video Moment Retrieval (VMR) targets to retrieve the specific moment corresponding to a sentence query from an untrimmed video. Although recent works have made remarkable progress…

cs.CV2026

Rethinking Video-Language Model from the Language Input Perspective

Xiang Fang, Wanlong Fang, Changshuo Wang +2

Driven by the wave of large language models, Video-Language Models (VLMs) have become a significant yet challenging technology to bridge the gap between videos and texts. Although…

cs.CV2026

Rethinking Weakly-supervised Video Temporal Grounding From a Game Perspective

Xiang Fang, Zeyu Xiong, Wanlong Fang +7

This paper addresses the challenging task of weakly-supervised video temporal grounding. Existing approaches are generally based on the moment proposal selection framework that uti…

cs.CV2026

VCE: A zero-cost hallucination mitigation method of LVLMs via visual contrastive editing

Yanbin Huang, Yisen Li, Guiyao Tie +6

Large vision-language models (LVLMs) frequently suffer from Object Hallucination (OH), wherein they generate descriptions containing objects that are not actually present in the in…

cs.CV2024

A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends

Daizong Liu, Mingyu Yang, Xiaoye Qu +3

With the significant development of large models in recent years, Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities across a wide range of multimodal u…