activity
20242026
most citedPseudolabel guided pixels contrast for domain adaptive semantic segmentation

7 citations · 7 across the 7 of their papers we have counts for

collaborators

9 papers

cs.CV2026

DSE-VTG: Dual-Side Enhancement for Training-Free Video Temporal Grounding

Zhuo Cao, Bingqing Zhang, Sen Wang +1

Text-guided Video Temporal Grounding (VTG) aims to localize the relevant segments in an untrimmed video based on text queries, yet collecting dense temporal annotations and trainin…

cs.IR2026

ReCoVR: Closing the Loop in Interactive Composed Video Retrieval

Bingqing Zhang, Yi Zhang, Zhuo Cao +4

Composed video retrieval (CoVR) searches for target videos using a reference video and a modification text, but existing methods are restricted to a single interaction round and ca…

cs.IR2026

Robust Test-time Video-Text Retrieval: Benchmarking and Adapting for Query Shifts

Bingqing Zhang, Zhuo Cao, Heming Du +4

Modern video-text retrieval (VTR) models excel on in-distribution benchmarks but are highly vulnerable to real-world query shifts, where the distribution of query data deviates fro…

cs.AI2026

AgentSelect: Benchmark for Narrative Query-to-Agent Recommendation

Yunxiao Shi, Wujiang Xu, Tingwei Chen +7

LLM agents are rapidly becoming the practical interface for task automation, yet the ecosystem lacks a principled way to choose among an exploding space of deployable configuration…

cs.CV2025

When One Moment Isn't Enough: Multi-Moment Retrieval with Cross-Moment Interactions

Zhuo Cao, Heming Du, Bingqing Zhang +3

Existing Moment retrieval (MR) methods focus on Single-Moment Retrieval (SMR). However, one query can correspond to multiple relevant moments in real-world applications. This makes…

cs.CV2025

Quantifying and Narrowing the Unknown: Interactive Text-to-Video Retrieval via Uncertainty Minimization

Bingqing Zhang, Zhuo Cao, Heming Du +4

Despite recent advances, Text-to-video retrieval (TVR) is still hindered by multiple inherent uncertainties, such as ambiguous textual queries, indistinct text-video mappings, and…