5 papers
Beyond Surface Imitation: Contrastive Modeling for Reasoning Path Alignment in Multimodal In-Context Learning
Mingbo Yang, Wenqiang Wang, Zhaolu Kang +4
In-context learning (ICL) is widely used in multimodal large language models (MLLMs) and achieves strong performance across a wide range of multimodal tasks. However, existing mult…
COMET: Contrastive Motion-Enhanced Temporal Reasoning for Video Multimodal Large Language Models
Chenghua Zhu, Zhaolu Kang, Qifan Shi +8
Video multimodal large language models have advanced significantly, yet fine-grained motion-temporal understanding remains fragile. The core bottleneck is not only sparse frame sam…
VFAD: Variational Semantic Prompting Meets Frequency-Adaptive Representation Learning for Zero-Shot Anomaly Detection
Peng Chen, Kaige Li, Wei Wang +5
Zero-shot anomaly detection (ZSAD) aims to detect and localize anomalies in unseen categories without access to target-specific training data. Although recent CLIP-based methods ha…
Domain Adaptive Object Detection via Dual-Stream Bilevel-Cycle Optimization
Yannan Chen, Wei Wang, Wenqiang Wang +5
Cycle self-training (CST) breaks the shared classifier assumption of the standard self-training framework, which is effective for unsupervised domain adaptation and exploits unlabe…
Towards Explainable Industrial Anomaly Detection via Knowledge-Guided Latent Reasoning
Peng Chen, Chao Huang, Yunkang Cao +7
Industrial anomaly detection demands precise reasoning over fine-grained defect patterns. However, existing multimodal large language models (MLLMs), pretrained on general-domain d…