collaborators

8 papers

cs.CV2026

Can We Perform Online RL for Image Editing without Editing Rewards?

Qichao Ma, Jikang Cheng, Ling Liang +3

Reinforcement learning (RL) enables direct preference optimization for image editing through editing-specific rewards, which remain less developed due to costly triplet supervision…

cs.CV2026

VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System

Haodong Li, Tianfei Ren, Xiaoxiao Ma +25

Text-to-video models have achieved remarkable visual quality, yet they still struggle to generate physically consistent dynamics because the temporal evolution of a scene must be i…

cs.RO2026

GenTrack: Physical Alignment for Robot-Native Motion Generation and Zero-Shot Humanoid Tracking

Zeyu Ling, Xinyao Yu, Renye Yan +4

General-purpose humanoid trackers can execute diverse references, but their zero-shot coverage depends on large embodied corpora that are costly to extend. Text-to-motion generator…

cs.LG2024

AdaMemento: Adaptive Memory-Assisted Policy Optimization for Reinforcement Learning

Renye Yan, Yaozhong Gan, You Wu +4

In sparse reward scenarios of reinforcement learning (RL), the memory mechanism provides promising shortcuts to policy optimization by reflecting on past experiences like humans. H…

cs.CL2024

Inner-Probe: Discovering Copyright-related Data Generation in LLM Architecture

Qichao Ma, Rui-Jie Zhu, Peiye Liu +8

Large Language Models (LLMs) utilize extensive knowledge databases and show powerful text generation ability. However, their reliance on high-quality copyrighted datasets raises co…

cs.LG2024

The Exploration-Exploitation Dilemma Revisited: An Entropy Perspective

Renye Yan, Yaozhong Gan, You Wu +4

The imbalance of exploration and exploitation has long been a significant challenge in reinforcement learning. In policy optimization, excessive reliance on exploration reduces lea…