Publications (29)
PrLM: Learning Explicit Reasoning for Personalized RAG via Contrastive Reward Optimization
Kepu Zhang, Teng Shi, Weijie Yu +1
Personalized retrieval-augmented generation (RAG) aims to produce user-tailored responses by incorporating retrieved user profiles alongside the input query. Existing methods prima…
Wasserstein Distance Regularized Sequence Representation for Text Matching in Asymmetrical Domains
Weijie Yu, Chen Xu, Jun Xu +4
One approach to matching texts from asymmetrical domains is projecting the input sequences into a common semantic space as feature vectors upon which the matching function can be r…
Hidden States Know Where Reasoning Diverges: Credit Assignment via Span-Level Wasserstein Distance
Xinzhu Chen, Wei He, Huichuan Fan +7
Group Relative Policy Optimization (GRPO) performs coarse-grained credit assignment in reinforcement learning with verifiable rewards (RLVR) by assigning the same advantage to all…
Similarity = Value? Consultation Value Assessment and Alignment for Personalized Search
Weicong Qin, Yi Xu, Weijie Yu +6
Personalized search systems in e-commerce platforms increasingly involve user interactions with AI assistants, where users consult about products, usage scenarios, and more. Levera…
An Explicit Syllogistic Legal Reasoning Framework for Large Language Models
Kepu Zhang, Weijie Yu, Zhongxiang Sun +1
Syllogistic reasoning is crucial for sound legal decision-making, allowing legal professionals to draw logical conclusions by applying general principles to specific case facts. Wh…
Legal Mathematical Reasoning with LLMs: Procedural Alignment through Two-Stage Reinforcement Learning
Kepu Zhang, Guofu Xie, Weijie Yu +4
Legal mathematical reasoning is essential for applying large language models (LLMs) in high-stakes legal contexts, where outputs must be both mathematically accurate and procedural…
CitaLaw: Enhancing LLM with Citations in Legal Domain
Kepu Zhang, Weijie Yu, Sunhao Dai +1
In this paper, we propose CitaLaw, the first benchmark designed to evaluate LLMs' ability to produce legally sound responses with appropriate citations. CitaLaw features a diverse…
Paragon: Parameter Generation for Controllable Multi-Task Recommendation
Chenglei Shen, Jiahao Zhao, Xiao Zhang +3
Commercial recommender systems face the challenge that task requirements from platforms or users often change dynamically (e.g., varying preferences for accuracy or diversity). Ide…
MoRE: A Mixture of Reflectors Framework for Large Language Model-Based Sequential Recommendation
Weicong Qin, Yi Xu, Weijie Yu +5
Large language models (LLMs) have emerged as a cutting-edge approach in sequential recommendation, leveraging historical interactions to model dynamic user preferences. Current met…
ReDeEP: Detecting Hallucination in Retrieval-Augmented Generation via Mechanistic Interpretability
Zhongxiang Sun, Xiaoxue Zang, Kai Zheng +5
Retrieval-Augmented Generation (RAG) models are designed to incorporate external knowledge, reducing hallucinations caused by insufficient parametric (internal) knowledge. However,…
ReARTeR: Retrieval-Augmented Reasoning with Trustworthy Process Rewarding
Zhongxiang Sun, Qipeng Wang, Weijie Yu +6
Retrieval-Augmented Generation (RAG) systems for Large Language Models (LLMs) hold promise in knowledge-intensive tasks but face limitations in complex multi-step reasoning. While…
GenRecEdit: Adapting Model Editing for Generative Recommendation with Cold-Start Items
Chenglei Shen, Teng Shi, Weijie Yu +2
Generative recommendation (GR) has shown strong potential for sequential recommendation in an end-to-end generation paradigm. However, existing GR models suffer from severe cold-st…
QE-RAG: A Robust Retrieval-Augmented Generation Benchmark for Query Entry Errors
Kepu Zhang, Zhongxiang Sun, Weijie Yu +5
Retriever-augmented generation (RAG) has become a widely adopted approach for enhancing the factual accuracy of large language models (LLMs). While current benchmarks evaluate the…
Logic Rules as Explanations for Legal Case Retrieval
Zhongxiang Sun, Kepu Zhang, Weijie Yu +2
In this paper, we address the issue of using logic rules to explain the results from legal case retrieval. The task is critical to legal case retrieval because the users (e.g., law…
Beyond Guilt: Legal Judgment Prediction with Trichotomous Reasoning
Kepu Zhang, Haoyue Yang, Xu Tang +2
In legal practice, judges apply the trichotomous dogmatics of criminal law, sequentially assessing the elements of the offense, unlawfulness, and culpability to determine whether a…
Decoding Recommendation Behaviors of In-Context Learning LLMs Through Gradient Descent
Yi Xu, Weicong Qin, Weijie Yu +3
Recently, there has been a growing trend in utilizing large language models (LLMs) for recommender systems, referred to as LLMRec. A notable approach within this trend is not to fi…
UOEP: User-Oriented Exploration Policy for Enhancing Long-Term User Experiences in Recommender Systems
Changshuo Zhang, Sirui Chen, Xiao Zhang +3
Reinforcement learning (RL) has gained traction for enhancing user long-term experiences in recommender systems by effectively exploring users' interests. However, modern recommend…
Uncovering ChatGPT's Capabilities in Recommender Systems
Sunhao Dai, Ninglu Shao, Haiyuan Zhao +6
The debut of ChatGPT has recently attracted the attention of the natural language processing (NLP) community and beyond. Existing studies have demonstrated that ChatGPT shows signi…
Explicitly Integrating Judgment Prediction with Legal Document Retrieval: A Law-Guided Generative Approach
Weicong Qin, Zelin Cao, Weijie Yu +3
Legal document retrieval and judgment prediction are crucial tasks in intelligent legal systems. In practice, determining whether two documents share the same judgments is essentia…
Deep Search with Hierarchical Meta-Cognitive Monitoring Inspired by Cognitive Neuroscience
Zhongxiang Sun, Qipeng Wang, Weijie Yu +3
Deep search agents powered by large language models have demonstrated strong capabilities in multi-step retrieval, reasoning, and long-horizon task execution. However, their practi…
Benefit from Rich: Tackling Search Interaction Sparsity in Search Enhanced Recommendation
Teng Shi, Weijie Yu, Xiao Zhang +3
In modern online platforms, search and recommendation (S&R) often coexist, offering opportunities for performance improvement through search-enhanced approaches. Existing studies s…
MAPS: Motivation-Aware Personalized Search via LLM-Driven Consultation Alignment
Weicong Qin, Yi Xu, Weijie Yu +5
Personalized product search aims to retrieve and rank items that match users' preferences and search intent. Despite their effectiveness, existing approaches typically assume that…
When Personalization Misleads: Understanding and Mitigating Hallucinations in Personalized LLMs
Zhongxiang Sun, Yi Zhan, Chenglei Shen +4
Personalized large language models (LLMs) adapt model behavior to individual users to enhance user satisfaction, yet personalization can inadvertently distort factual reasoning. We…
Towards Understanding Continual Factual Knowledge Acquisition of Language Models: From Theory to Algorithm
Haoyu Wang, Yifan Shang, Zhongxiang Sun +3
Continual Pre-Training (CPT) is essential for enabling Language Models (LMs) to integrate new knowledge without erasing old. While classical CPT techniques like data replay have be…
LLaDA-Rec: Discrete Diffusion for Parallel Semantic ID Generation in Generative Recommendation
Teng Shi, Chenglei Shen, Weijie Yu +6
Generative recommendation represents each item as a semantic ID, i.e., a sequence of discrete tokens, and generates the next item through autoregressive decoding. While effective,…
Bridging Search and Recommendation through Latent Cross Reasoning
Teng Shi, Weicong Qin, Weijie Yu +4
Search and recommendation (S&R) are fundamental components of modern online platforms, yet effectively leveraging search behaviors to improve recommendation remains a challenging p…
Beyond High-Entropy Exploration: Correctness-Aware Low-Entropy Segment-Based Advantage Shaping for Reasoning LLMs
Xinzhu Chen, Xuesheng Li, Zhongxiang Sun +1
Reinforcement Learning with Verifiable Rewards (RLVR) has become a central approach for improving the reasoning ability of large language models. Recent work studies RLVR through t…
Enhancing Bandit Algorithms with LLMs for Time-varying User Preferences in Streaming Recommendations
Chenglei Shen, Yi Zhan, Weijie Yu +2
In real-world streaming recommender systems, user preferences evolve dynamically over time. Existing bandit-based methods treat time merely as a timestamp, neglecting its explicit…
Explainable Legal Case Matching via Inverse Optimal Transport-based Rationale Extraction
Weijie Yu, Zhongxiang Sun, Jun Xu +4
As an essential operation of legal retrieval, legal case matching plays a central role in intelligent legal systems. This task has a high demand on the explainability of matching r…