3 papers
cs.LG2026
Granularity-Adaptive Credit Assignment for Long-Horizon LLM Agent Reinforcement Learning
Taoran Liang, Yang Liu, Shang Luo +9
Reinforcement learning is now the standard way to train large language model agents on long-horizon tasks, where dozens of interdependent actions precede a single sparse reward. Cr…
cs.AI2026
Harness-agnostic detection and immunization of reward hacking in self-evolving language models
Rongxin Yang, Yang Liu, Shang Luo +10
Self-evolving language models improve by proposing candidate updates and keeping whatever raises a visible score. When that score is an imperfect proxy for the capability one actua…
cs.AI2026
DiffImaginE: Imagine to Verify Entity Types with Diffusion
Feng Zhang, Feiyu Han, Rongxin Yang +11
Multimodal named entity recognition (MNER) determines whether each candidate span and entity-type hypothesis is supported by joint textual and visual evidence. Existing imagine-and…