collaborators

8 papers

cs.CV2026

MirrorPPR: Exemplar-Based Portrait Photo Retouching

Zhihong Liu, Zheng Li, Jiachun Jin +4

While text-guided image editing has made remarkable progress, it remains limited in structural portrait retouching. Textual descriptions struggle to convey fine-grained changes to…

cs.RO2026

World-Language-Action Model for Unified World Modeling, Language Reasoning, and Action Synthesis

Yi Yang, Zhihong Liu, Siqi Kou +9

We propose world-language-action (WLA) models as a new class of embodied foundation models. WLA takes textual instructions, images, and robot states as inputs to jointly predict te…

cs.CV2026

ProductWebGen: Benchmarking Multimodal Product Webpage Generation

Zhihong Liu, Siqi Kou, Zheng Li +5

Crafting a product display webpage from a source product image, along with layout and visual content instructions, holds significant practical value for domains such as marketing,…

cs.CV2026

LatentUM: Unleashing the Potential of Interleaved Cross-Modal Reasoning via a Latent-Space Unified Model

Jiachun Jin, Zetong Zhou, Xiao Yang +4

Unified models (UMs) hold promise for their ability to understand and generate content across heterogeneous modalities. Compared to merely generating visual content, the use of UMs…

cs.CV2026

Mantis: A Versatile Vision-Language-Action Model with Disentangled Visual Foresight

Yi Yang, Xueqi Li, Yiyang Chen +7

Recent advances in Vision-Language-Action (VLA) models demonstrate that visual signals can effectively complement sparse action supervisions. However, letting VLA directly predict…

cs.CV2025

LiveTalk: Real-Time Multimodal Interactive Video Diffusion via Improved On-Policy Distillation

Ethan Chern, Zhulin Hu, Bohao Tang +4

Real-time video generation via diffusion is essential for building general-purpose multimodal interactive AI systems. However, the simultaneous denoising of all video frames with b…