From the 1 of 75 linked papers with an AI index.
1 citations · 1 across the 27 of their papers we have counts for
6 papers · 1 filter
Different Feedback, Different Updates: Selective Self-Learning from User Interactions for Large Language Models
Xuanchen Li, Haitao Li, Yujia Zhou +5
User feedback offers natural supervision for persistent LLM improvement, but a single message may support multiple behavioral changes with different scopes of generalization. We in…
MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop
Xuancheng Li, Haitao Li, Yujia Zhou +2
Reinforcement Learning with Verifiable Rewards (RLVR) is widely used to improve reasoning across domains, but outcome-only scalar rewards are often sparse and uninformative. This l…
How do Humans Process AI-generated Hallucination Contents: a Neuroimaging Study
Shuqi Zhu, Yi Zhong, Ziyi Ye +4
While AI-generated hallucinations pose considerable risks, the underlying cognitive mechanisms by which humans can successfully recognize or be misled by these hallucinations remai…
From <Answer> to <Think>: Multidimensional Supervision of Reasoning Process for LLM Optimization
Beining Wang, Weihang Su, Hongtao Tian +5
Improving the multi-step reasoning ability of Large Language Models (LLMs) is a critical yet challenging task. The dominant paradigm, outcome-supervised reinforcement learning (RLV…
JustEva: A Toolkit to Evaluate LLM Fairness in Legal Knowledge Inference
Zongyue Xue, Siyuan Zheng, Shaochun Wang +8
The integration of Large Language Models (LLMs) into legal practice raises pressing concerns about judicial fairness, particularly due to the nature of their "black-box" processes.…
Evaluating Intelligence via Trial and Error
Jingtao Zhan, Jiahao Zhao, Jiayu Li +7
Intelligence is a crucial trait for species to find solutions within a limited number of trial-and-error attempts. Building on this idea, we introduce Survival Game as a framework…