13 citations · 28 across the 18 of their papers we have counts for
20 papers · 1 filter
Let ViT Speak: Generative Language-Image Pre-training
Yan Fang, Mengcheng Lan, Zilong Huang +7
In this paper, we present \textbf{Gen}erative \textbf{L}anguage-\textbf{I}mage \textbf{P}re-training (GenLIP), a minimalist generative pretraining framework for Vision Transformers…
AlignGen: Boosting Personalized Image Generation with Cross-Modality Prior Alignment
Yiheng Lin, Shifang Zhao, Ting Liu +4
Personalized image generation aims to integrate user-provided concepts into text-to-image models, enabling the generation of customized content based on a given prompt. Recent zero…
PreFM: Online Audio-Visual Event Parsing via Predictive Future Modeling
Xiao Yu, Yan Fang, Xiaojie Jin +2
Audio-visual event parsing plays a crucial role in understanding multimodal video content, but existing methods typically rely on offline processing of entire videos with huge mode…
OmniAD: Detect and Understand Industrial Anomaly via Multimodal Reasoning
Shifang Zhao, Yiheng Lin, Lu Han +2
While anomaly detection has made significant progress, generating detailed analyses that incorporate industrial knowledge remains a challenge. To address this gap, we introduce Omn…
3DResT: A Strong Baseline for Semi-Supervised 3D Referring Expression Segmentation
Wenxin Chen, Mengxue Qu, Weitai Kang +3
3D Referring Expression Segmentation (3D-RES) typically requires extensive instance-level annotations, which are time-consuming and costly. Semi-supervised learning (SSL) mitigates…
IPSeg: Image Posterior Mitigates Semantic Drift in Class-Incremental Segmentation
Xiao Yu, Yan Fang, Yao Zhao +1
Class incremental learning aims to enable models to learn from sequential, non-stationary data streams across different tasks without catastrophic forgetting. In class incremental…