activity
20242026
collaborators

18 papers

cs.LG2026

Your LLM, Your Style: Behavioral Mode Axes for LLM Behavioral Control

Haoze Liu, Run Liu, Haiying Xu +6

Large language models (LLMs) increasingly act in interactive settings where their behavioral styles affect user experience, safety, and downstream decision making. Existing LLM per…

cs.AI2026

: An End-to-End Agent Auditing Engine

Haoning Wang, Mingxun Zhang, Chenyue Yu +4

With the rapid advancement of large language models (LLMs), harnesses have become essential infrastructure for deploying agents across a wide range of domains. The fast-evolving ha…

cs.AI2026

SkillEval: Decomposing Agent Skill Quality into Interpretable Signals

Jiahui Han, Qinuo Li, Ziheng Peng +6

Agent skills provide reusable procedural knowledge that helps agents solve specialized tasks. As their use expands, evaluating skill quality becomes increasingly important. Existin…

cs.CL2026

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning

Xiuyi Lou, Zicheng Xu, Yu-Neng Chuang +4

Reinforcement learning (RL) has achieved remarkable success in enhancing the reasoning capabilities of large language models (LLMs). However, widely used critic-free RL methods rel…

cs.CL2026

DynamicMem: A Long-Horizon Memory Benchmark in Real-World Settings

Wenya Xie, Shengming Zhou, Zelin Li +9

LLM agents increasingly act as personal assistants that must remember a user's profile over months: who they are (attributes), what they routinely do (habits), and what they prefer…

cs.CL2026

Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning

Zicheng Xu, Ruixuan Zhang, Yu-Neng Chuang +7

Large Language Models (LLMs) achieve remarkable reasoning capabilities through reinforcement learning (RL) post-training. However, existing RL post-training commonly relies on unif…