11 papers
Multimodal Reward Hacking in Reinforcement Learning
Jiayu Yao, Yiwei Wang, Anmeng Zhang +5
Reinforcement learning (RL) is increasingly used to align multimodal large language models (MLLMs), but higher rewards do not always imply better task performance. This risk is amp…
SAFARI: Scaling Long Horizon Agentic Fault Attribution via Active Investigation
Chenyang Zhu, Jiayu Yao, Kushal Chawla +10
As autonomous agents tackle increasingly complex multi-step, multi-agent tasks, their execution trajectories have scaled beyond the constraints of even the largest context windows.…
HighlightBench: Benchmarking Markup-Driven Table Reasoning in Scientific Documents
Lexin Wang, Shenghua Liu, Yiwei Wang +5
Visual markups such as highlights, underlines, and bold text are common in table-centric documents. Although multimodal large language models (MLLMs) have made substantial progress…
PRISM-: Differential Subspace Steering for Prompt Highlighting in Large Language Models
Yuyao Ge, Shenghua Liu, Yiwei Wang +6
Prompt highlighting steers a large language model to prioritize user-specified text spans during generation. A key challenge of existing Key-editing approaches is extracting steeri…
AudioMotionBench: Evaluating Auditory Motion Perception in Audio LLMs
Zhe Sun, Yujun Cai, Jiayu Yao +1
Large Audio-Language Models (LALMs) have recently shown impressive progress in speech recognition, audio captioning, and auditory question answering. Yet, whether these models can…
Gated Differentiable Working Memory for Long-Context Language Modeling
Lingrui Mei, Shenghua Liu, Yiwei Wang +7
Long contexts challenge transformers: attention scores dilute across thousands of tokens, critical information is often lost in the middle, and models struggle to adapt to novel pa…