8 papers · 1 filter
Multimodal Reward Hacking in Reinforcement Learning
Jiayu Yao, Yiwei Wang, Anmeng Zhang +5
Reinforcement learning (RL) is increasingly used to align multimodal large language models (MLLMs), but higher rewards do not always imply better task performance. This risk is amp…
Supervised Fine-tuning with Synthetic Rationale Data Hurts Real-World Disease Prediction
Buxin Su, Bingxuan Li, Cheng Qian +3
Supervised fine-tuning with synthetic rationale data is widely assumed to improve language model performance on clinical prediction tasks by teaching models not just what to predic…
Harmonizing Dense and Sparse Signals in Multi-turn RL: Dual-Horizon Credit Assignment for Industrial Sales Agents
Haojin Yang, Ai Jian, Xinyue Huang +5
Optimizing large language models for industrial sales requires balancing long-term commercial objectives (e.g., conversion rate) with immediate linguistic constraints such as fluen…
PromptCD: Test-Time Behavior Enhancement via Polarity-Prompt Contrastive Decoding
Baolong Bi, Yuyao Ge, Shenghua Liu +9
Reliable AI systems require large language models (LLMs) to exhibit behaviors aligned with human preferences and values. However, most existing alignment approaches operate at trai…
A Survey of Vibe Coding with Large Language Models
Yuyao Ge, Lingrui Mei, Zenghao Duan +12
The advancement of large language models (LLMs) has catalyzed a paradigm shift from code generation assistance to autonomous coding agents, enabling a novel development methodology…
Reward and Guidance through Rubrics: Promoting Exploration to Improve Multi-Domain Reasoning
Baolong Bi, Shenghua Liu, Yiwei Wang +6
Recent advances in reinforcement learning (RL) have significantly improved the complex reasoning capabilities of large language models (LLMs). Despite these successes, existing met…