4 papers · 1 filter
H2HMem: A Multimodal Memory Benchmark for Agents in Human-Human Interactions
Shiping Zhu, Yibo Yang, Zhengyang Wang +3
Large language model agents are increasingly deployed in human-human interaction settings, such as meeting assistants and clinical documentation systems, where they must observe co…
Blockwise SFT for Diffusion Language Models: Reconciling Bidirectional Attention and Autoregressive Decoding
Bowen Sun, Yujun Cai, Ming-Hsuan Yang +1
Discrete diffusion language models have shown strong potential for text generation, yet standard supervised fine-tuning (SFT) misaligns with their semi-autoregressive inference: tr…
Structured Attention Matters to Multimodal LLMs in Document Understanding
Chang Liu, Hongkai Chen, Yujun Cai +4
Document understanding remains a significant challenge for multimodal large language models (MLLMs). While previous research has primarily focused on locating evidence pages throug…
MIRAGE: Multimodal Immersive Reasoning and Guided Exploration for Red-Team Jailbreak Attacks
Wenhao You, Bryan Hooi, Yiwei Wang +5
While safety mechanisms have significantly progressed in filtering harmful text inputs, MLLMs remain vulnerable to multimodal jailbreaks that exploit their cross-modal reasoning ca…