13 papers
Diagnosing Corruption-Induced Reliability Failures in Vision-Language Models
Xiangjie Sui, Songyang Li, Hanwei Zhu +3
Visual corruptions can change vision--language model (VLM) behavior in ways that top-1 accuracy does not capture. A model may keep the same answer while losing distributional suppo…
ChronoLock: Protecting Videos from Unauthorized Text-to-Video Personalization
Jiaming He, Jiashu Zhang, Guanyu Hou +4
Text-to-video (T2V) diffusion models have made it increasingly easy to synthesize realistic and temporally coherent videos, while recent personalization techniques allow such model…
EduVQA: Towards Concept-Aware Assessment of Educational AI-Generated Videos
Baoliang Chen, Xinlong Bu, Hanwei Zhu +2
Existing AI-generated video quality assessment (AIGVQA) methods mainly focus on global perceptual realism and coarse text-video alignment, while overlooking a critical requirement…
Q-Tacit: Image Quality Assessment via Latent Visual Reasoning
Yuxuan Jiang, Yixuan Li, Hanwei Zhu +3
Vision-Language Model (VLM)-based image quality assessment (IQA) has been significantly advanced by incorporating Chain-of-Thought (CoT) reasoning. Recent work has refined image qu…
Beyond Cosine Similarity: Magnitude-Aware CLIP for No-Reference Image Quality Assessment
Zhicheng Liao, Dongxu Wu, Zhenshan Shi +5
Recent efforts have repurposed the Contrastive Language-Image Pre-training (CLIP) model for No-Reference Image Quality Assessment (NR-IQA) by measuring the cosine similarity betwee…
Simple Lines, Big Ideas: Towards Interpretable Assessment of Human Creativity from Drawings
Zihao Lin, Zhenshan Shi, Sasa Zhao +4
Assessing human creativity through visual outputs, such as drawings, plays a critical role in fields including psychology, education, and cognitive science. However, current assess…