6 citations · 10 across the 7 of their papers we have counts for
7 papers
Instance Brownian Bridge as Texts for Open-vocabulary Video Instance Segmentation
Zesen Cheng, Kehan Li, Hao Li +5
Temporally locating objects with arbitrary class texts is the primary pursuit of open-vocabulary Video Instance Segmentation (VIS). Because of the insufficient vocabulary of video…
AutoConv: Automatically Generating Information-seeking Conversations with Large Language Models
Siheng Li, Cheng Yang, Yichun Yin +6
Information-seeking conversation, which aims to help users gather information through conversation, has achieved great progress in recent years. However, the research is still stym…
NewsDialogues: Towards Proactive News Grounded Conversation
Siheng Li, Yichun Yin, Cheng Yang +7
Hot news is one of the most popular topics in daily conversations. However, news grounded conversation has long been stymied by the lack of well-designed task definition and scarce…
Text-Video Retrieval with Disentangled Conceptualization and Set-to-Set Alignment
Peng Jin, Hao Li, Zesen Cheng +5
Text-video retrieval is a challenging cross-modal task, which aims to align visual entities with natural language descriptions. Current methods either fail to leverage the local de…
TG-VQA: Ternary Game of Video Question Answering
Hao Li, Peng Jin, Zesen Cheng +5
Video question answering aims at answering a question about the video content by reasoning the alignment semantics within them. However, since relying heavily on human instructions…
Multi-granularity Interaction Simulation for Unsupervised Interactive Segmentation
Kehan Li, Yian Zhao, Zhennan Wang +6
Interactive segmentation enables users to segment as needed by providing cues of objects, which introduces human-computer interaction for many fields, such as image editing and med…