activity
20232026
most citedBeyond Visual Cues: Synchronously Exploring Target-Centric Semantics for Vision-Language Tracking

3 citations · 3 across the 9 of their papers we have counts for

collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV2026

ViewSAM: Learning View-aware Cross-modal Semantics for Weakly Supervised Cross-view Referring Multi-Object Tracking

Jiawei Ge, Xintian Zhang, Jiuxin Cao +9

Cross-view Referring Multi-Object Tracking (CRMOT) aims to track multiple objects specified by natural language across multiple camera views, with globally consistent identities. D…

cs.CV2025

Debate-Enhanced Pseudo Labeling and Frequency-Aware Progressive Debiasing for Weakly-Supervised Camouflaged Object Detection with Scribble Annotations

Jiawei Ge, Jiuxin Cao, Xinyi Li +5

Weakly-Supervised Camouflaged Object Detection (WSCOD) aims to locate and segment objects that are visually concealed within their surrounding scenes, relying solely on sparse supe…

cs.CV2025

Denoise-then-Retrieve: Text-Conditioned Video Denoising for Video Moment Retrieval

Weijia Liu, Jiuxin Cao, Bo Miao +6

Current text-driven Video Moment Retrieval (VMR) methods encode all video clips, including irrelevant ones, disrupting multimodal alignment and hindering optimization. To this end,…

cs.CV2024

Context-Enhanced Video Moment Retrieval with Large Language Models

Weijia Liu, Bo Miao, Jiuxin Cao +4

Current methods for Video Moment Retrieval (VMR) struggle to align complex situations involving specific environmental details, character descriptions, and action narratives. To ta…

cs.CV2024

Query-Based Knowledge Sharing for Open-Vocabulary Multi-Label Classification

Xuelin Zhu, Jian Liu, Dongqi Tang +4

Identifying labels that did not appear during training, known as multi-label zero-shot learning, is a non-trivial task in computer vision. To this end, recent studies have attempte…

cs.CV2023

Text as Image: Learning Transferable Adapter for Multi-Label Classification

Xuelin Zhu, Jiuxin Cao, Jian liu +7

Pre-trained vision-language models have notably accelerated progress of open-world concept recognition. Their impressive zero-shot ability has recently been transferred to multi-la…