10 papers
Agentic Reinforcement Learning with Self-Distilled Reward Shaping
Ranxu Zhang, Guinan Chen, Chenshaodong +5
Agentic reinforcement learning enables LLM agents to learn through interaction, but sparse trajectory-level rewards reveal success without identifying which intermediate decisions…
MCLMR: A Model-Agnostic Causal Learning Framework for Multi-Behavior Recommendation
Ranxu Zhang, Junjie Meng, Ying Sun +5
Multi-Behavior Recommendation (MBR) leverages multiple user interaction types (e.g., views, clicks, purchases) to enrich preference modeling and alleviate data sparsity issues in t…
From Correctness to Preference: A Framework for Personalized Agentic Reinforcement Learning
Ranxu zhang, zeyang li, Jiacheng Huang +5
Agentic reinforcement learning (Agentic RL) has achieved strong progress in tasks with clear success signals. However, many real-world agent applications require user-conditioned b…
VERDICT: Verifiable Evolving Reasoning with Directive-Informed Collegial Teams for Legal Judgment Prediction
Hui Liao, Chuan Qin, Yongwen Ren +4
Legal Judgment Prediction (LJP) predicts applicable law articles, charges, and penalty terms from case facts. Beyond accuracy, LJP calls for intrinsically interpretable and legally…
Latent Shadows: The Gaussian-Discrete Duality in Masked Diffusion
Guinan Chen, Xunpeng Huang, Ying Sun +3
Masked discrete diffusion is a dominant paradigm for high-quality language modeling where tokens are iteratively corrupted to a mask state, yet its inference efficiency is bottlene…
Rethinking Popularity Bias in Collaborative Filtering via Analytical Vector Decomposition
Lingfeng Liu, Yixin Song, Dazhong Shen +4
Popularity bias fundamentally undermines the personalization capabilities of collaborative filtering (CF) models, causing them to disproportionately recommend popular items while n…