5 papers
See Only When Needed: Context-Aware Attention Intervention for Mitigating Hallucinations in LVLMs
Yuqing Lei, Wenbo Lyu, Yingjun Du +3
Large Vision-Language Models (LVLMs) excel at multimodal tasks but remain prone to object hallucinations. Prior training-free remedies often uniformly strengthen visual signals, wh…
A Physically-Grounded Attack and Adaptive Defense Framework for Real-World Low-Light Image Enhancement
Tongshun Zhang, Pingping Liu, Yuqing Lei +3
Limited illumination often causes severe physical noise and detail degradation in images. Existing Low-Light Image Enhancement (LLIE) methods frequently treat the enhancement proce…
Personalized Safety Alignment for Text-to-Image Diffusion Models
Yu Lei, Jinbin Bai, Qingyu Shi +4
Text-to-image diffusion models have revolutionized visual content generation, yet their deployment is hindered by a fundamental limitation: safety mechanisms enforce rigid, uniform…
MetaTPT: Meta Test-time Prompt Tuning for Vision-Language Models
Yuqing Lei, Yingjun Du, Yawen Huang +2
Vision-language models (VLMs) such as CLIP exhibit strong zero-shot generalization but remain sensitive to domain shifts at test time. Test-time prompt tuning (TPT) mitigates this…
From Masks to Worlds: A Hitchhiker's Guide to World Models
Jinbin Bai, Yu Lei, Hecong Wu +7
This is not a typical survey of world models; it is a guide for those who want to build worlds. We do not aim to catalog every paper that has ever mentioned a ``world model". Inste…