4 papers
SafeSieve: From Heuristics to Experience in Progressive Pruning for LLM-based Multi-Agent Communication
Ruijia Zhang, Xinyan Zhao, Ruixiang Wang +5
LLM-based multi-agent systems exhibit strong collaborative capabilities but often suffer from redundant communication and excessive token overhead. Existing methods typically enhan…
POLAR: A Pessimistic Model-based Policy Learning Algorithm for Dynamic Treatment Regimes
Ruijia Zhang, Xiangyu Zhang, Zhengling Qi +2
Dynamic treatment regimes (DTRs) provide a principled framework for optimizing sequential decision-making in domains where decisions must adapt over time in response to individual…
Improved Rates of Differentially Private Nonconvex-Strongly-Concave Minimax Optimization
Ruijia Zhang, Mingxi Lei, Meng Ding +3
In this paper, we study the problem of (finite sum) minimax optimization in the Differential Privacy (DP) model. Unlike most of the previous studies on the (strongly) convex-concav…
Understanding Inverse Reinforcement Learning under Overparameterization: Non-Asymptotic Analysis and Global Optimality
Ruijia Zhang, Siliang Zeng, Chenliang Li +2
The goal of the Inverse reinforcement learning (IRL) task is to identify the underlying reward function and the corresponding optimal policy from a set of expert demonstrations. Wh…