collaborators

8 papers

cs.CV2025

CtrlVDiff: Controllable Video Generation via Unified Multimodal Video Diffusion

Dianbing Xi, Jiepeng Wang, Yuanzhi Liang +8

We tackle the dual challenges of video understanding and controllable video generation within a unified diffusion framework. Our key insights are two-fold: geometry-only cues (e.g.…

cs.CV2025

Learning What to Trust: Bayesian Prior-Guided Optimization for Visual Generation

Ruiying Liu, Yuanzhi Liang, Haibin Huang +2

Group Relative Policy Optimization (GRPO) has emerged as an effective and lightweight framework for post-training visual generative models. However, its performance is fundamentall…

cs.CV2025

UniModel: A Visual-Only Framework for Unified Multimodal Understanding and Generation

Chi Zhang, Jiepeng Wang, Youming Wang +5

We present UniModel, a unified generative model that jointly supports visual understanding and visual generation within a single pixel-to-pixel diffusion framework. Our goal is to…

cs.AI2025

ScRPO: From Errors to Insights

Lianrui Li, Dakuan Lu, Jiawei Shao +1

We introduce Self-correction Relative Policy Optimization (ScRPO), a novel reinforcement learning framework designed to empower large language models with advanced mathematical rea…

cs.LG2025

CAS-Spec: Cascade Adaptive Self-Speculative Decoding for On-the-Fly Lossless Inference Acceleration of LLMs

Zhiyuan Ning, Jiawei Shao, Ruge Xu +4

Speculative decoding has become a widely adopted as an effective technique for lossless inference acceleration when deploying large language models (LLMs). While on-the-fly self-sp…

cs.CL2025

Pipeline Parallelism is All You Need for Optimized Early-Exit Based Self-Speculative Decoding

Ruanjun Li, Ziheng Liu, Yuanming Shi +3

Large language models (LLMs) deliver impressive generation quality, but incur very high inference cost because each output token is generated auto-regressively through all model la…