3 papers
cs.AI2026
TRACER: Per-Tool Context Retention for LLM Agents via Consequence-Attributed Reinforcement Learning
Ziqi Lin, Ye Wu, Mengying Yang +4
Enterprise data agents answer business queries by chaining many tool calls over multiple reasoning steps, routinely accumulating hundreds of thousands of context tokens per session…
cs.IR2026
DREAM Technical Report
Bin Zhang, Bowen Zheng, Chao Yi +74
Industrial recommender systems commonly use cascaded retrieval, ranking, and re-ranking pipelines. Although efficient, these pipelines fragment information and objectives across mo…
cs.IR2026
Objective Shaping with Hard Negatives: Windowed Partial AUC Optimization for RL-based LLM Recommenders
Wentao Shi, Qifan Wang, Chen Chen +7
Reinforcement learning (RL) effectively optimizes Large Language Model (LLM)-based recommenders by contrasting positive and negative items. Empirically, training with beam-search n…