activity
20242026
collaborators

13 papers

cs.CV2026

CLIP-GS: Unifying Vision-Language Representation with 3D Gaussian Splatting

Siyu Jiao, Haoye Dong, Yuyang Yin +5

Recent works in 3D multimodal learning have made remarkable progress. However, typically 3D multimodal models are only capable of handling point clouds. Compared to the emerging 3D…

cs.CV2025

PreFM: Online Audio-Visual Event Parsing via Predictive Future Modeling

Xiao Yu, Yan Fang, Xiaojie Jin +2

Audio-visual event parsing plays a crucial role in understanding multimodal video content, but existing methods typically rely on offline processing of entire videos with huge mode…

cs.CV2025

OmniAD: Detect and Understand Industrial Anomaly via Multimodal Reasoning

Shifang Zhao, Yiheng Lin, Lu Han +2

While anomaly detection has made significant progress, generating detailed analyses that incorporate industrial knowledge remains a challenge. To address this gap, we introduce Omn…

cs.CV2025

AlignGen: Boosting Personalized Image Generation with Cross-Modality Prior Alignment

Yiheng Lin, Shifang Zhao, Ting Liu +4

Personalized image generation aims to integrate user-provided concepts into text-to-image models, enabling the generation of customized content based on a given prompt. Recent zero…

cs.CV2025

3DResT: A Strong Baseline for Semi-Supervised 3D Referring Expression Segmentation

Wenxin Chen, Mengxue Qu, Weitai Kang +3

3D Referring Expression Segmentation (3D-RES) typically requires extensive instance-level annotations, which are time-consuming and costly. Semi-supervised learning (SSL) mitigates…

cs.CV2025

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model

Kunyang Han, Yibo Hu, Mengxue Qu +3

Advances in CLIP and large multimodal models (LMMs) have enabled open-vocabulary and free-text segmentation, yet existing models still require predefined category prompts, limiting…