12 papers
Exposure Bias Can Alleviate Itself via Directional and Frequency Rectification in Flow Matching
Guanbo Huang, Jingjia Mao, Fanding Huang +9
Flow Matching (FM) has achieved remarkable generative performance, yet it suffers from exposure bias due to discrepancies between training and inference. Existing mitigation strate…
Physics-Driven Spatiotemporal Modeling for AI-Generated Video Detection
Shuhai Zhang, ZiHao Lian, Jiahao Yang +6
AI-generated videos have achieved near-perfect visual realism (e.g., Sora), urgently necessitating reliable detection mechanisms. However, detecting such videos faces significant c…
Multi-Turn Adaptive Prompting Attack on Large Vision-Language Models
In Chong Choi, Jiacheng Zhang, Feng Liu +1
Multi-turn jailbreak attacks have proven effective against text-only large language models (LLMs), where malicious content is gradually introduced to bypass safety alignment. Howev…
SAVAA: Mitigating Hallucinations in LVLMs via Step-wise Adaptive Visual Attention Amplification
Jiacheng Zhang, Feng Liu, Chao Du +1
A line of recent training-free methods for mitigating hallucinations in large vision-language models (LVLMs) operates by amplifying attention to visual tokens during autoregressive…
ReflexFlow: Rethinking Learning Objective for Exposure Bias Alleviation in Flow Matching
Guanbo Huang, Jingjia Mao, Fanding Huang +8
Despite tremendous recent progress, Flow Matching methods still suffer from exposure bias due to discrepancies in training and inference. This paper investigates the root causes of…
Vision Language Models Cannot Plan, but Can They Formalize?
Muyu He, Yuxi Zheng, Yuchen Liu +7
The advancement of vision language models (VLMs) has empowered embodied agents to accomplish simple multimodal planning tasks, but not long-horizon ones requiring long sequences of…