From the 1 of 5 linked papers with an AI index.
5 papers
SVR-R1: Bootstrapping Multi-modal Reasoning with Self-verification in Reinforcement Learning
Mingyuan Wu, Jingcheng Yang, Shengyi Qian +11
The paper introduces SVR-R1, a reinforcement learning framework that lets a multimodal model generate an answer and then self‑verify it with a binary verdict, allowing a second‑cha…
Objective Shaping with Hard Negatives: Windowed Partial AUC Optimization for RL-based LLM Recommenders
Wentao Shi, Qifan Wang, Chen Chen +7
Reinforcement learning (RL) effectively optimizes Large Language Model (LLM)-based recommenders by contrasting positive and negative items. Empirically, training with beam-search n…
Training-Free Test-Time Contrastive Learning for Large Language Models
Kaiwen Zheng, Kai Zhou, Jinwu Hu +3
Large language models (LLMs) demonstrate strong reasoning capabilities, but their performance often degrades under distribution shift. Existing test-time adaptation (TTA) methods r…
Verifiable Reasoning for LLM-based Generative Recommendation
Xinyu Lin, Hanqing Zeng, Hanchao Yu +8
Reasoning in Large Language Models (LLMs) has recently shown strong potential in enhancing generative recommendation through deep understanding of complex user preference. Existing…
REMEDI: Relative Feature Enhanced Meta-Learning with Distillation for Imbalanced Prediction
Fei Liu, Huanhuan Ren, Yu Guan +4
Predicting future vehicle purchases among existing owners presents a critical challenge due to extreme class imbalance (<0.5% positive rate) and complex behavioral patterns. We pro…