3 papers
cs.CV2026
From Hallucination to Grounding: Diagnosing Visual Spatial Intelligence via CRISP
Zhixing Li, Yinan Yu
Current VLM evaluations often conflate language priors with genuine spatial reasoning. To address this, we introduce CRISP, a novel structural-diagnostic evaluation paradigm that a…
cs.LG2026
Latent Domain Prompt Learning for Vision-Language Models
Zhixing Li, Arsham Gholamzadeh Khoee, Yinan Yu
The objective of domain generalization (DG) is to enable models to be robust against domain shift. DG is crucial for deploying vision-language models (VLMs) in real-world applicati…
cs.CR2024
JailPO: A Novel Black-box Jailbreak Framework via Preference Optimization against Aligned LLMs
Hongyi Li, Jiawei Ye, Jie Wu +3
Large Language Models (LLMs) aligned with human feedback have recently garnered significant attention. However, it remains vulnerable to jailbreak attacks, where adversaries manipu…