7 papers
SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback
Qianxi Yan, Chunrong Chen, Jiuzhou Zhao +3
Agent Skills are today either hand-authored or produced in a single LLM generation pass, and consequently possess no closed loop through which they might improve from the interacti…
SPOT: Sparse Probing and Outcome Calibration for On-Policy Distillation
Zikun Qu, Min Zhang, Mingze Kong +5
On-policy distillation (OPD) provides dense teacher supervision on student-generated trajectories, but standard reverse-KL training can assign insufficient probability to other pla…
T-POP: Test-Time Personalization with Online Preference Feedback
Zikun Qu, Min Zhang, Mingze Kong +7
Personalizing large language models (LLMs) to individual user preferences is a critical step beyond generating generically helpful responses. However, current personalization metho…
ALSO: Adversarial Online Strategy Optimization for Social Agents
Xiang Li, Liping Yi, Mingze Kong +3
Social simulation provides a compelling testbed for studying social intelligence, where agents interact through multi-turn dialogues under evolving contexts and strategically adapt…
TCAndon-Router: Adaptive Reasoning Router for Multi-Agent Collaboration
Jiuzhou Zhao, Chunrong Chen, Chenqi Qiao +3
Multi-Agent Systems(MAS) have become a powerful paradigm for building high performance intelligent applications. Within these systems, the router responsible for determining which…
Generative Sequential Recommendation via Hierarchical Behavior Modeling
Zhefan Wang, Guokai Yan, Jinbei Yu +5
Recommender systems in multi-behavior domains, such as advertising and e-commerce, aim to guide users toward high-value but inherently sparse conversions. Leveraging auxiliary beha…