3 papers
cs.CV2025
MetaTPT: Meta Test-time Prompt Tuning for Vision-Language Models
Yuqing Lei, Yingjun Du, Yawen Huang +2
Vision-language models (VLMs) such as CLIP exhibit strong zero-shot generalization but remain sensitive to domain shifts at test time. Test-time prompt tuning (TPT) mitigates this…
cs.LG2025
From Masks to Worlds: A Hitchhiker's Guide to World Models
Jinbin Bai, Yu Lei, Hecong Wu +7
This is not a typical survey of world models; it is a guide for those who want to build worlds. We do not aim to catalog every paper that has ever mentioned a ``world model". Inste…
cs.CV2025
Personalized Safety Alignment for Text-to-Image Diffusion Models
Yu Lei, Jinbin Bai, Qingyu Shi +4
Text-to-image diffusion models have revolutionized visual content generation, yet their deployment is hindered by a fundamental limitation: safety mechanisms enforce rigid, uniform…