1 citations · 1 across the 7 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Different Feedback, Different Updates: Selective Self-Learning from User Interactions for Large Language Models
Xuanchen Li, Haitao Li, Yujia Zhou +5
User feedback offers natural supervision for persistent LLM improvement, but a single message may support multiple behavioral changes with different scopes of generalization. We in…
cs.AI2026
MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop
Xuancheng Li, Haitao Li, Yujia Zhou +2
Reinforcement Learning with Verifiable Rewards (RLVR) is widely used to improve reasoning across domains, but outcome-only scalar rewards are often sparse and uninformative. This l…