6 papers
Different Feedback, Different Updates: Selective Self-Learning from User Interactions for Large Language Models
Xuanchen Li, Haitao Li, Yujia Zhou +5
User feedback offers natural supervision for persistent LLM improvement, but a single message may support multiple behavioral changes with different scopes of generalization. We in…
Adaptive Multi-Resolution Procedural Knowledge Compression for Large Language Models
Changyue Wang, Weihang Su, Qingyao Ai +5
Large language models (LLMs) are widely used to tackle complex tasks with autonomous workflows. Recently, reusable natural language skills have emerged as a popular paradigm to inj…
MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop
Xuancheng Li, Haitao Li, Yujia Zhou +2
Reinforcement Learning with Verifiable Rewards (RLVR) is widely used to improve reasoning across domains, but outcome-only scalar rewards are often sparse and uninformative. This l…
Beyond Exposure: Optimizing Ranking Fairness with Non-linear Time-Income Functions
Xuancheng Li, Tao Yang, Yujia Zhou +2
Ranking systems in web search and recommendation allocate attention among items and providers, and therefore need to balance relevance-based effectiveness with provider fairness. E…
Beyond Experience Retrieval: Learning to Generate Utility-Optimized Structured Experience for Frozen LLMs
Xuancheng Li, Haitao Li, Yujia Zhou +2
Large language models (LLMs) are largely static and often redo reasoning or repeat mistakes. Prior experience reuse typically relies on external retrieval, which is similarity-base…
ATACompressor: Adaptive Task-Aware Compression for Efficient Long-Context Processing in LLMs
Xuancheng Li, Haitao Li, Yujia Zhou +2
Long-context inputs in large language models (LLMs) often suffer from the "lost in the middle" problem, where critical information becomes diluted or ignored due to excessive lengt…