50 papers
Learning from Unreachable Rewards: Hint-Conditioned Reinforcement Learning for Generative Recommendation
Kangning Zhang, Haotian Fang, Xukun Luo +6
Semantic-ID generative recommenders represent each item as a short sequence of discrete semantic tokens and predict the next item by autoregressively generating this token sequence…
Fast, Slow, and Tool-augmented Thinking for LLMs: A Review
Xinda Jia, Jinpeng Li, Zezhong Wang +6
Large Language Models (LLMs) have demonstrated remarkable progress in reasoning across diverse domains. However, effective reasoning in real-world tasks requires adapting the reaso…
BALTO: Balanced Token-Level Policy Optimization for Hallucination Mitigation
Ning Li, Zixuan Guo, Yan Xu +7
Hallucinations remain a major obstacle to deploying large language models (LLMs) in knowledge-intensive settings, where generated responses must be faithfully grounded in provided…
DiffCold: A Diffusion-based Generative Model for Cold-Start Item Recommendation
Kangning Zhang, Yingjie Qin, Weinan Zhang +2
Cold-start item recommendation remains a persistent challenge in real-world systems due to the absence of interaction histories. While prior models attempt to bridge this gap using…
SkillJuror: Measuring How Agent Skill Organization Changes Runtime Behavior
Zhiyu Chen, Zihan Guo, Bo Huang +4
Agent Skills augment large language model (LLM) agents with procedural knowledge at inference time, but current benchmarks rarely distinguish what a Skill says from how it is organ…
MOTOR: Learning ID-free Item Representation with Token Crossing for Embedding-based Multimodal Recommendation
Kangning Zhang, Jiarui Jin, Yingjie Qin +4
While multimodal recommendation models have effectively integrated visual and textual information, their reliance on unique ID embeddings constitutes a fundamental performance bott…