2 papers
cs.CV2025
Neuro-Symbolic Spatial Reasoning in Segmentation
Jiayi Lin, Jiabo Huang, Shaogang Gong
Open-Vocabulary Semantic Segmentation (OVSS) assigns pixel-level labels from an open set of categories, requiring generalization to unseen and unlabelled objects. Using vision-lang…
cs.CV2024
MLLM as Video Narrator: Mitigating Modality Imbalance in Video Moment Retrieval
Weitong Cai, Jiabo Huang, Shaogang Gong +2
Video Moment Retrieval (VMR) aims to localize a specific temporal segment within an untrimmed long video given a natural language query. Existing methods often suffer from inadequa…