14 papers
ExMesh: EXplicit Mesh Reconstruction with Topology Adaptation
Chuanjin Fan, Lifan Wu, Wenjie Chang +3
Reconstructing surface meshes from multi-view images has remained a core challenge in recent years. Most existing methods, whether implicit or explicit, depend on intermediate repr…
Adversarial Attacks Already Tell the Answer: Directional Bias-Guided Test-time Defense for Vision-Language Models
Liangsheng Liu, Si Chen, Jiamin Wu +5
Vision-Language Models (VLMs), such as CLIP, have shown strong zero-shot generalization but remain highly vulnerable to adversarial perturbations, posing serious risks in real-worl…
ComPose: A Unified Completion-Pose Framework for Robust Category-Level Object Pose Estimation
Huan Ren, Yihan Chen, Chuxin Wang +3
Category-level object pose estimation aims to predict the pose and size of arbitrary objects in specific categories. Existing methods struggle with the inherent incompleteness of o…
GeoGuide: Hierarchical Geometric Guidance for Open-Vocabulary 3D Semantic Segmentation
Xujing Tao, Chuxin Wang, Yubo Ai +8
Open-vocabulary 3D semantic segmentation aims to segment arbitrary categories beyond the training set. Existing methods predominantly rely on distilling knowledge from 2D open-voca…
BindWeave: Subject-Consistent Video Generation via Cross-Modal Integration
Zhaoyang Li, Dongjun Qian, Kai Su +6
Diffusion Transformer has shown remarkable abilities in generating high-fidelity videos, delivering visually coherent frames and rich details over extended durations. However, exis…
SMTrack: State-Aware Mamba for Efficient Temporal Modeling in Visual Tracking
Yinchao Ma, Dengqing Yang, Zhangyu He +2
Visual tracking aims to automatically estimate the state of a target object in a video sequence, which is challenging especially in dynamic scenarios. Thus, numerous methods are pr…