7 papers
CREST: Curvature-Regulated Event-Centric Sampling for Efficient Long-Video Understanding
Mehrajul Abadin Miraj, Abdul Mohaimen Al Radi, Shariful Islam Rayhan +4
Selecting informative frames from long videos is a combinatorial problem that existing methods address either through efficient heuristics without explicit modeling of query-condit…
SafeSteer: A Decoding-level Defense Mechanism for Multimodal Large Language Models
Xinyi Zeng, Xue Yang, Jingyuan Zhang +5
Multimodal large language models (MLLMs) are gaining increasing attention. Due to the heterogeneity of their input features, they face significant challenges in terms of jailbreak…
Do Models See in Line with Human Vision? Probing the Correspondence Between LVLM Representations and EEG Signals
Xin Xiao, Yang Lei, Haoyang Zeng +6
Large Vision Language Models (LVLMs) exhibit strong visual understanding and reasoning abilities. However, whether their internal representations reflect human visual cognition is…
Do Models Hear Like Us? Probing the Representational Alignment of Audio LLMs and Naturalistic EEG
Haoyun Yang, Xin Xiao, Jiang Zhong +5
Audio Large Language Models (Audio LLMs) have demonstrated strong capabilities in integrating speech perception with language understanding. However, whether their internal represe…
ES-Mem: Event Segmentation-Based Memory for Long-Term Dialogue Agents
Huhai Zou, Tianhao Sun, Chuanjiang He +6
Memory is critical for dialogue agents to maintain coherence and enable continuous adaptation in long-term interactions. While existing memory mechanisms offer basic storage and re…
DiffER: Diffusion Entity-Relation Modeling for Reversal Curse in Diffusion Large Language Models
Shaokai He, Kaiwen Wei, Xinyi Zeng +5
The "reversal curse" refers to the phenomenon where large language models (LLMs) exhibit predominantly unidirectional behavior when processing logically bidirectional relationships…