6 papers
Geometry-Preserving Orthonormal Initialization for Low-Rank Adaptation in RLVR
Ruijia Zhang, Jiacheng Zhu, Hanqing Zhu +1
Low-rank adaptation (LoRA) and its variants enable parameter-efficient fine-tuning of large language models under the supervised fine-tuning (SFT) paradigm. However, their efficacy…
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization
Yihang Yao, Zhepeng Cen, Haohong Lin +6
Proactive large language model (LLM) agents aim to actively plan, query, and interact over multiple turns, enabling efficient task completion beyond passive instruction following a…
MAS-PromptBench: When Does Prompt Optimization Improve Multi-Agent LLM Systems?
Juyang Bai, Laixi Shi
Multi-agent systems (MAS) offer a scalable path forward for agentic AI, comprising multiple LLM-based agents, each assigned a system prompt and a position within a workflow that go…
Taming the Curses of Multiagency in Robust Markov Games with Large State Space through Linear Function Approximation
Jingchu Gai, Laixi Shi
Multi-agent reinforcement learning (MARL) holds great potential but faces robustness challenges due to environmental uncertainty. To address this, distributionally robust Markov ga…
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning
Yiqing Liang, Jielin Qiu, Wenhao Ding +7
Reinforcement Learning with Verifiable Rewards (RLVR) has recently emerged as a powerful paradigm for post-training large language models (LLMs), achieving state-of-the-art perform…
BECAUSE: Bilinear Causal Representation for Generalizable Offline Model-based Reinforcement Learning
Haohong Lin, Wenhao Ding, Jian Chen +4
Offline model-based reinforcement learning (MBRL) enhances data efficiency by utilizing pre-collected datasets to learn models and policies, especially in scenarios where explorati…