17 papers
Mitigating Identity Essentialism in LLM Agents with Longitudinal Life Trajectories
Hexi Wang, Yujia Zhou, Bangde Du +7
Large language models (LLMs) offer a scalable approach to social simulation, but their credibility depends on how agents are constructed. Existing methods can partially reproduce p…
Different Feedback, Different Updates: Selective Self-Learning from User Interactions for Large Language Models
Xuanchen Li, Haitao Li, Yujia Zhou +5
User feedback offers natural supervision for persistent LLM improvement, but a single message may support multiple behavioral changes with different scopes of generalization. We in…
Structure-aware Relative Policy Optimization for Ranking
Yiteng Tu, Weihang Su, Zitao Su +3
Ranking is a fundamental component of modern information access systems. Reinforcement learning (RL) provides a flexible framework for directly optimizing coarse-grained feedback a…
Co-Evolving LLM Evaluators and Policies via DynamicRubric
Beining Wang, Weihang Su, Hongtao Tian +8
Post-training with evaluator feedback on policy-induced samples serves as a major mechanism for improving large language models. As policies improve, these sampled responses become…
Generative Chinese Statute Retrieval
Yiteng Tu, Zitao Su, Weihang Su +5
The paper introduces GCSR, a generative framework that treats Chinese statute retrieval as a sequence generation task and embeds hierarchical legal knowledge into the model to impr…
MetaSyn: A Benchmark for LLM Agents on Meta-Analysis Articles from Nature Portfolio
Anzhe Xie, Weihang Su, Yujia Zhou +3
Systematic review and meta-analysis is an important method for scientific research. It comprehensively studies target research questions by combining evidence from multiple indepen…