5 papers
What to Remove, What to Preserve: Dual-Ambiguity Rectification for All-in-One Image Restoration
Cencen Liu, Wen Yin, Dongyang Zhang +6
The paper introduces DAR-Net, a deep network that tackles the dual ambiguity problem in all‑in‑one image restoration by modeling degradation states with a simplex‑constrained arche…
CBV: Clean-label Backdoor Attacks on Vision Language Models via Diffusion Models
Ji Guo, Xiaolong Qin, Cencen Liu +3
Vision-Language Models (VLMs) have achieved remarkable success in tasks such as image captioning and visual question answering (VQA). However, as their applications become increasi…
One Model, Two Minds: Task-Conditioned Reasoning for Unified Image Quality and Aesthetic Assessment
Wen Yin, Cencen Liu, Dingrui Liu +3
Unifying Image Quality Assessment (IQA) and Image Aesthetic Assessment (IAA) in a single multimodal large language model is appealing, yet existing methods adopt a task-agnostic re…
AlignVAR: Towards Globally Consistent Visual Autoregression for Image Super-Resolution
Cencen Liu, Dongyang Zhang, Wen Yin +6
Visual autoregressive (VAR) models have recently emerged as a promising alternative for image generation, offering stable training, non-iterative inference, and high-fidelity synth…
TiCAL:Typicality-Based Consistency-Aware Learning for Multimodal Emotion Recognition
Wen Yin, Siyu Zhan, Cencen Liu +5
Multimodal Emotion Recognition (MER) aims to accurately identify human emotional states by integrating heterogeneous modalities such as visual, auditory, and textual data. Existing…