11 papers
ProductWebGen: Benchmarking Multimodal Product Webpage Generation
Zhihong Liu, Siqi Kou, Zheng Li +5
Crafting a product display webpage from a source product image, along with layout and visual content instructions, holds significant practical value for domains such as marketing,…
Generative Recommendation for Large-Scale Advertising
Ben Xue, Dan Liu, Lixiang Wang +27
Generative recommendation has recently attracted widespread attention in industry due to its potential for scaling and stronger model capacity. However, deploying real-time generat…
Adaptive Video Distillation: Mitigating Oversaturation and Temporal Collapse in Few-Step Generation
Yuyang You, Yongzhi Li, Jiahui Li +3
Video generation has recently emerged as a central task in the field of generative AI. However, the substantial computational cost inherent in video synthesis makes model distillat…
Think-Then-Generate: Reasoning-Aware Text-to-Image Diffusion with LLM Encoders
Siqi Kou, Jiachun Jin, Zetong Zhou +8
Recent progress in text-to-image (T2I) diffusion models (DMs) has enabled high-quality visual synthesis from diverse textual prompts. Yet, most existing T2I DMs, even those equippe…
EEA: Exploration-Exploitation Agent for Long Video Understanding
Te Yang, Xiangyu Zhu, Bo Wang +3
Long-form video understanding requires efficient navigation of extensive visual data to pinpoint sparse yet critical information. Current approaches to longform video understanding…
Music Grounding by Short Video
Zijie Xin, Minquan Wang, Jingyu Liu +4
Adding proper background music helps complete a short video to be shared. Previous work tackles the task by video-to-music retrieval (V2MR), aiming to find the most suitable music…