17 papers
Multimodal Reward Hacking in Reinforcement Learning
Jiayu Yao, Yiwei Wang, Anmeng Zhang +5
Reinforcement learning (RL) is increasingly used to align multimodal large language models (MLLMs), but higher rewards do not always imply better task performance. This risk is amp…
SAW: Stage-Aware Dynamic Weighting for Multi-Objective Reinforcement Learning in Large Language Models
Yuchen He, Baolong Bi, Shenghua Liu +7
Although multi-objective reinforcement learning (MORL) is central to aligning large language models with complex human preferences, the prevailing practice of static weighted summa…
SkillAttack: Automated Red Teaming of Agent Skills through Attack Path Refinement
Zenghao Duan, Yuxin Tian, Zhiyi Yin +6
LLM-based agent systems increasingly rely on agent skills sourced from open registries to extend their capabilities, yet the openness of such ecosystems makes skills difficult to t…
HighlightBench: Benchmarking Markup-Driven Table Reasoning in Scientific Documents
Lexin Wang, Shenghua Liu, Yiwei Wang +5
Visual markups such as highlights, underlines, and bold text are common in table-centric documents. Although multimodal large language models (MLLMs) have made substantial progress…
PRISM-: Differential Subspace Steering for Prompt Highlighting in Large Language Models
Yuyao Ge, Shenghua Liu, Yiwei Wang +6
Prompt highlighting steers a large language model to prioritize user-specified text spans during generation. A key challenge of existing Key-editing approaches is extracting steeri…
PromptCD: Test-Time Behavior Enhancement via Polarity-Prompt Contrastive Decoding
Baolong Bi, Yuyao Ge, Shenghua Liu +9
Reliable AI systems require large language models (LLMs) to exhibit behaviors aligned with human preferences and values. However, most existing alignment approaches operate at trai…