works on

From the 1 of 9 linked papers with an AI index.

collaborators

9 papers

cs.CV2026

Sample-Adaptive Latent Rewards for Uncertainty-Guided Diffusion Post-Training

Rui Li, Yuanzhi Liang, Ke Hao +4

Latent reward models can supervise visual diffusion models without decoding intermediate states into pixel space. This makes alignment with human preferences more efficient. Howeve…

cs.CV2026

Rethinking Reward Signals in Video GRPO: When Scores Become Targets

Rui Li, Yuanzhi Liang, Ziqi Ni +3

The paper proposes TaRoS, a framework that redesigns reward signals for video generation using GRPO to avoid reward hacking and saturation, improving visual fidelity, motion cohere…

cs.CV2026

Seeing What Matters: Visual Preference Policy Optimization for Visual Generation

Ziqi Ni, Yuanzhi Liang, Rui Li +4

Reinforcement learning (RL) has become a powerful tool for post-training visual generative models, with Group Relative Policy Optimization (GRPO) increasingly used to align generat…

cs.CV2026

Learning to Credit the Right Steps: Objective-aware Process Optimization for Visual Generation

Rui Li, Ke Hao, Yuanzhi Liang +4

Reinforcement learning, particularly Group Relative Policy Optimization (GRPO), has emerged as an effective framework for post-training visual generative models with human preferen…

cs.CV2026

Reward-Aware Trajectory Shaping for Few-step Visual Generation

Rui Li, Bingyu Li, Yuanzhi Liang +3

Achieving high-fidelity generation in extremely few sampling steps has long been a central goal of generative modeling. Existing approaches largely rely on distillation-based frame…

cs.CL2026

Stochasticity in Tokenisation Improves Robustness

Sophie Steger, Rui Li, Sofiane Ennadir +4

The widespread adoption of large language models (LLMs) has increased concerns about their robustness. Vulnerabilities in perturbations of tokenisation of the input indicate that m…