10 papers
MirrorPPR: Exemplar-Based Portrait Photo Retouching
Zhihong Liu, Zheng Li, Jiachun Jin +4
While text-guided image editing has made remarkable progress, it remains limited in structural portrait retouching. Textual descriptions struggle to convey fine-grained changes to…
World-Language-Action Model for Unified World Modeling, Language Reasoning, and Action Synthesis
Yi Yang, Zhihong Liu, Siqi Kou +9
We propose world-language-action (WLA) models as a new class of embodied foundation models. WLA takes textual instructions, images, and robot states as inputs to jointly predict te…
ProductWebGen: Benchmarking Multimodal Product Webpage Generation
Zhihong Liu, Siqi Kou, Zheng Li +5
Crafting a product display webpage from a source product image, along with layout and visual content instructions, holds significant practical value for domains such as marketing,…
Think-Then-Generate: Reasoning-Aware Text-to-Image Diffusion with LLM Encoders
Siqi Kou, Jiachun Jin, Zetong Zhou +8
Recent progress in text-to-image (T2I) diffusion models (DMs) has enabled high-quality visual synthesis from diverse textual prompts. Yet, most existing T2I DMs, even those equippe…
Fast and Accurate Causal Parallel Decoding using Jacobi Forcing
Lanxiang Hu, Siqi Kou, Yichao Fu +5
Multi-token generation has emerged as a promising paradigm for accelerating transformer-based large model inference. Recent efforts primarily explore diffusion Large Language Model…
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks
Yi Yang, Jiaxuan Sun, Siqi Kou +2
Real-world embodied agents face long-horizon tasks, characterized by high-level goals demanding multi-step solutions beyond single actions. Successfully navigating these requires b…