most citedUniSOT: A Unified Framework for Multi-Modality Single Object Tracking

4 citations · 5 across the 8 of their papers we have counts for

collaborators

11 papers

cs.CV2026

PriorPose: Reference-Guided Joint Deformation and Alignment for Category-Level Object Pose Estimation

Yihan Chen, Huan Ren, Wenfei Yang +3

Category-level object pose estimation seeks to recover a similarity transform for unseen instances without instance-specific CAD models. Most competitive methods are corr…

cs.CV2026

Adversarial Attacks Already Tell the Answer: Directional Bias-Guided Test-time Defense for Vision-Language Models

Liangsheng Liu, Si Chen, Jiamin Wu +5

Vision-Language Models (VLMs), such as CLIP, have shown strong zero-shot generalization but remain highly vulnerable to adversarial perturbations, posing serious risks in real-worl…

cs.CV2026

ComPose: A Unified Completion-Pose Framework for Robust Category-Level Object Pose Estimation

Huan Ren, Yihan Chen, Chuxin Wang +3

Category-level object pose estimation aims to predict the pose and size of arbitrary objects in specific categories. Existing methods struggle with the inherent incompleteness of o…

cs.CV20254 cited

UniSOT: A Unified Framework for Multi-Modality Single Object Tracking

Yinchao Ma, Yuyang Tang, Wenfei Yang +3

Single object tracking aims to localize target object with specific reference modalities (bounding box, natural language or both) in a sequence of specific video modalities (RGB, R…

cs.CV2025

BindWeave: Subject-Consistent Video Generation via Cross-Modal Integration

Zhaoyang Li, Dongjun Qian, Kai Su +6

Diffusion Transformer has shown remarkable abilities in generating high-fidelity videos, delivering visually coherent frames and rich details over extended durations. However, exis…

cs.CV2025

StruMamba3D: Exploring Structural Mamba for Self-supervised Point Cloud Representation Learning

Chuxin Wang, Yixin Zha, Wenfei Yang +1

Recently, Mamba-based methods have demonstrated impressive performance in point cloud representation learning by leveraging State Space Model (SSM) with the efficient context model…