2 papers
cs.CV2025
LLaVA-KD: A Framework of Distilling Multimodal Large Language Models
Yuxuan Cai, Jiangning Zhang, Haoyang He +7
The success of Large Language Models (LLMs) has inspired the development of Multimodal Large Language Models (MLLMs) for unified understanding of vision and language. However, the…
cs.CV2025
Omni-AD: Learning to Reconstruct Global and Local Features for Multi-class Anomaly Detection
Jiajie Quan, Ao Tong, Yuxuan Cai +3
In multi-class unsupervised anomaly detection(MUAD), reconstruction-based methods learn to map input images to normal patterns to identify anomalous pixels. However, this strategy…