works on

From the 1 of 7 linked papers with an AI index.

collaborators

7 papers

cs.RO2026

RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance

Dongchi Huang, Hongyin Zhang, Bohan Hou +12

General-purpose reward models are increasingly the bottleneck for scaling robot learning, yet the recipe for learning value-related capabilities from large-scale heterogeneous corp…

cs.LG2026

TAPO: Transition-Aware Policy Optimization for LLM Agents

Cong Li, Peixi Peng, Yisen Zhao +4

The paper introduces TAPO, a training framework that augments reinforcement learning for large language model agents with action‑conditioned next‑observation prediction, improving…

cs.CV2026

LiViBench: An Omnimodal Benchmark for Interactive Livestream Video Understanding

Xiaodong Wang, Langling Huang, Zhirong Wu +4

The development of multimodal large language models (MLLMs) has advanced general video understanding. However, existing video evaluation benchmarks primarily focus on non-interacti…

cs.CV2026

COVR:Collaborative Optimization of VLMs and RL Agent for Visual-Based Control

Canming Xia, Peixi Peng, Guang Tan +4

Visual reinforcement learning (RL) suffers from poor sample efficiency due to high-dimensional observations in complex tasks. While existing works have shown that vision-language m…

cs.CV2025

LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs

Xiaodong Wang, Jinfa Huang, Li Yuan +1

Most Video Large Language Models (Video-LLMs) adopt preference alignment techniques, e.g., DPO~\citep{rafailov2024dpo}, to optimize the reward margin between a winning response ($y…

cs.CV2025

LongDWM: Cross-Granularity Distillation for Building a Long-Term Driving World Model

Xiaodong Wang, Zhirong Wu, Peixi Peng

Driving world models are used to simulate futures by video generation based on the condition of the current state and actions. However, current models often suffer serious error ac…