collaborators

5 papers

cs.CV2026

Multiple Consistent 2D-3D Mappings for Robust Zero-Shot 3D Visual Grounding

Yufei Yin, Jie Zheng, Qianke Meng +7

Zero-shot 3D Visual Grounding (3DVG) is a critical capability for open-world embodied AI. However, existing methods are fundamentally bottlenecked by the poor quality of open-vocab…

cs.CV2026

Progressive Video Condensation with MLLM Agent for Long-form Video Understanding

Yufei Yin, Yuchen Xing, Qianke Meng +3

Understanding long videos requires extracting query-relevant information from long sequences under tight compute budgets. Existing text-then-LLM pipelines lose fine-grained visual…

cs.CV2026

VideoARM: Agentic Reasoning over Hierarchical Memory for Long-Form Video Understanding

Yufei Yin, Qianke Meng, Minghao Chen +3

Long-form video understanding remains challenging due to the extended temporal structure and dense multimodal cues. Despite recent progress, many existing approaches still rely on…

cs.CV2026

Self-Classification Enhancement and Correction for Weakly Supervised Object Detection

Yufei Yin, Lechao Cheng, Wengang Zhou +3

In recent years, weakly supervised object detection (WSOD) has attracted much attention due to its low labeling cost. The success of recent WSOD models is often ascribed to the two…

cs.CV2026

HERO: Hierarchical Embedding-Refinement for Open-Vocabulary Temporal Sentence Grounding in Videos

Tingting Han, Xinsong Tao, Yufei Yin +3

Temporal Sentence Grounding in Videos (TSGV) aims to temporally localize segments of a video that correspond to a given natural language query. Despite recent progress, most existi…