collaborators

8 papers

cs.IR2026

Towards Fast Domain Adaptation and Fine-Grained User Simulation for Evaluating Conversational Recommender Systems

Yuanzi Li, Quanyu Dai, Xueyang Feng +5

Conversational Recommender Systems (CRSs) enhance user experience through multi-turn interactions, yet evaluating their performance remains challenging. While Large Language Model…

cs.AI2026

ARCO: Adaptive Rubrics with Co-Evolution for Multi-Step LLM-Based Agents

Zihang Tian, Jingsen Zhang, Rui Li +3

Reinforcement learning for multi-step LLM agents often relies on scalar rewards that indicate success but cannot explain why a trajectory is good or bad. Rubric-based rewards impro…

cs.IR2026

Do Generative Recommenders Deepen the Information Cocoon? A Closed-Loop Simulation with LLM-powered User Simulators

Jiyuan Yang, Gengxin Sun, Mengqi Zhang +5

Recommender systems alleviate information overload, yet repeated feedback between recommendations and user interactions can reinforce existing preferences and narrow users' exposur…

cs.CY2026

Benchmarking LLMs for Community Governance Simulation with Life-history Narratives

Xu Chen, Yuanzi Li, Lei Wang +6

Effective community governance hinges on understanding what specific residents think and need. Recent work has used large language models (LLMs) to simulate human respondents, offe…

cs.CL2026

Evaluating and Calibrating LLM Confidence on Questions with Multiple Correct Answers

Yuhan Wang, Shiyu Ni, Zhikai Ding +3

Confidence calibration is essential for making large language models (LLMs) reliable, yet existing training-free methods have been primarily studied under single-answer question an…

cs.CV2026

Seeing Before Agreeing: Aligning Multi-Agent Consensus with Visual Evidence

Yuhan Wang, Shuochen Chang, Yalin Feng +8

Vision-language models (VLMs) have achieved strong performance on visual question answering (VQA). To mitigate individual hallucinations and blind spots, aggregating diverse perspe…