works on

From the 1 of 64 linked papers with an AI index.

activity
20242026
most citedMetaSyn: A Benchmark for LLM Agents on Meta-Analysis Articles from Nature Portfolio

1 citations · 1 across the 7 of their papers we have counts for

collaborators

64 papers

cs.CL2026

Mitigating Identity Essentialism in LLM Agents with Longitudinal Life Trajectories

Hexi Wang, Yujia Zhou, Bangde Du +7

Large language models (LLMs) offer a scalable approach to social simulation, but their credibility depends on how agents are constructed. Existing methods can partially reproduce p…

cs.AI2026

Different Feedback, Different Updates: Selective Self-Learning from User Interactions for Large Language Models

Xuanchen Li, Haitao Li, Yujia Zhou +5

User feedback offers natural supervision for persistent LLM improvement, but a single message may support multiple behavioral changes with different scopes of generalization. We in…

cs.IR2026

Structure-aware Relative Policy Optimization for Ranking

Yiteng Tu, Weihang Su, Zitao Su +3

Ranking is a fundamental component of modern information access systems. Reinforcement learning (RL) provides a flexible framework for directly optimizing coarse-grained feedback a…

cs.LG2026

Co-Evolving LLM Evaluators and Policies via DynamicRubric

Beining Wang, Weihang Su, Hongtao Tian +8

Post-training with evaluator feedback on policy-induced samples serves as a major mechanism for improving large language models. As policies improve, these sampled responses become…

cs.IR2026

Generative Chinese Statute Retrieval

Yiteng Tu, Zitao Su, Weihang Su +5

The paper introduces GCSR, a generative framework that treats Chinese statute retrieval as a sequence generation task and embeds hierarchical legal knowledge into the model to impr…

cs.DL2026

RWGBench: Evaluating Scholarly Positioning in Related Work Generation

Anzhe Xie, Weihang Su, Jiaxin Mao +4

Large language models have shown strong fluency in scientific writing, yet the evaluation of related work generation (RWG) remains limited. Existing RWG evaluations largely inherit…