5 papers
ContextNav: Towards Agentic Multimodal In-Context Learning
Honghao Fu, Yuan Ouyang, Kai-Wei Chang +3
Recent advances demonstrate that multimodal large language models (MLLMs) exhibit strong multimodal in-context learning (ICL) capabilities, enabling them to adapt to novel vision-l…
FrameMind: Frame-Interleaved Video Reasoning via Reinforcement Learning
Haonan Ge, Yiwei Wang, Kai-Wei Chang +2
Current video understanding models rely on fixed frame sampling strategies, processing predetermined visual inputs regardless of the specific reasoning requirements of each questio…
DRS: Deep Question Reformulation With Structured Output
Zhecheng Li, Yiwei Wang, Bryan Hooi +3
Question answering represents a core capability of large language models (LLMs). However, when individuals encounter unfamiliar knowledge in texts, they often formulate questions t…
Enhancing LLM Character-Level Manipulation via Divide and Conquer
Zhen Xiong, Yujun Cai, Bryan Hooi +3
Large Language Models (LLMs) have demonstrated strong generalization capabilities across a wide range of natural language processing (NLP) tasks. However, they exhibit notable weak…
Exploring Visual Vulnerabilities via Multi-Loss Adversarial Search for Jailbreaking Vision-Language Models
Shuyang Hao, Bryan Hooi, Jun Liu +3
Despite inheriting security measures from underlying language models, Vision-Language Models (VLMs) may still be vulnerable to safety alignment issues. Through empirical analysis,…