activity
20242026
collaborators

10 papers

cs.LG2026

Agentic Reinforcement Learning with Self-Distilled Reward Shaping

Ranxu Zhang, Guinan Chen, Chenshaodong +5

Agentic reinforcement learning enables LLM agents to learn through interaction, but sparse trajectory-level rewards reveal success without identifying which intermediate decisions…

cs.IR2026

MCLMR: A Model-Agnostic Causal Learning Framework for Multi-Behavior Recommendation

Ranxu Zhang, Junjie Meng, Ying Sun +5

Multi-Behavior Recommendation (MBR) leverages multiple user interaction types (e.g., views, clicks, purchases) to enrich preference modeling and alleviate data sparsity issues in t…

cs.CL2026

From Correctness to Preference: A Framework for Personalized Agentic Reinforcement Learning

Ranxu zhang, zeyang li, Jiacheng Huang +5

Agentic reinforcement learning (Agentic RL) has achieved strong progress in tasks with clear success signals. However, many real-world agent applications require user-conditioned b…

cs.IR2026

VERDICT: Verifiable Evolving Reasoning with Directive-Informed Collegial Teams for Legal Judgment Prediction

Hui Liao, Chuan Qin, Yongwen Ren +4

Legal Judgment Prediction (LJP) predicts applicable law articles, charges, and penalty terms from case facts. Beyond accuracy, LJP calls for intrinsically interpretable and legally…

cs.LG2026

Latent Shadows: The Gaussian-Discrete Duality in Masked Diffusion

Guinan Chen, Xunpeng Huang, Ying Sun +3

Masked discrete diffusion is a dominant paradigm for high-quality language modeling where tokens are iteratively corrupted to a mask state, yet its inference efficiency is bottlene…

cs.IR2026

Rethinking Popularity Bias in Collaborative Filtering via Analytical Vector Decomposition

Lingfeng Liu, Yixin Song, Dazhong Shen +4

Popularity bias fundamentally undermines the personalization capabilities of collaborative filtering (CF) models, causing them to disproportionately recommend popular items while n…