3 papers
cs.AI2025
ContextNav: Towards Agentic Multimodal In-Context Learning
Honghao Fu, Yuan Ouyang, Kai-Wei Chang +3
Recent advances demonstrate that multimodal large language models (MLLMs) exhibit strong multimodal in-context learning (ICL) capabilities, enabling them to adapt to novel vision-l…
cs.AI2025
RefineShot: Rethinking Cinematography Understanding with Foundational Skill Evaluation
Hang Wu, Yujun Cai, Haonan Ge +3
Cinematography understanding refers to the ability to recognize not only the visual content of a scene but also the cinematic techniques that shape narrative meaning. This capabili…
cs.CV2024
Exploring Visual Vulnerabilities via Multi-Loss Adversarial Search for Jailbreaking Vision-Language Models
Shuyang Hao, Bryan Hooi, Jun Liu +3
Despite inheriting security measures from underlying language models, Vision-Language Models (VLMs) may still be vulnerable to safety alignment issues. Through empirical analysis,…