works on

From the 1 of 5 linked papers with an AI index.

collaborators

5 papers

cs.AI2026

SVR-R1: Bootstrapping Multi-modal Reasoning with Self-verification in Reinforcement Learning

Mingyuan Wu, Jingcheng Yang, Shengyi Qian +11

The paper introduces SVR-R1, a reinforcement learning framework that lets a multimodal model generate an answer and then self‑verify it with a binary verdict, allowing a second‑cha…

cs.IR2026

Objective Shaping with Hard Negatives: Windowed Partial AUC Optimization for RL-based LLM Recommenders

Wentao Shi, Qifan Wang, Chen Chen +7

Reinforcement learning (RL) effectively optimizes Large Language Model (LLM)-based recommenders by contrasting positive and negative items. Empirically, training with beam-search n…

cs.CL2026

Training-Free Test-Time Contrastive Learning for Large Language Models

Kaiwen Zheng, Kai Zhou, Jinwu Hu +3

Large language models (LLMs) demonstrate strong reasoning capabilities, but their performance often degrades under distribution shift. Existing test-time adaptation (TTA) methods r…

cs.IR2026

Verifiable Reasoning for LLM-based Generative Recommendation

Xinyu Lin, Hanqing Zeng, Hanchao Yu +8

Reasoning in Large Language Models (LLMs) has recently shown strong potential in enhancing generative recommendation through deep understanding of complex user preference. Existing…

cs.LG2025

REMEDI: Relative Feature Enhanced Meta-Learning with Distillation for Imbalanced Prediction

Fei Liu, Huanhuan Ren, Yu Guan +4

Predicting future vehicle purchases among existing owners presents a critical challenge due to extreme class imbalance (<0.5% positive rate) and complex behavioral patterns. We pro…