activity
20222026
collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV2026

Joint Alignment and Distillation for Video Generation via Sample-Guided Distribution Matching

Jiuzhou Lin, Junlong Wu, Fei Zuo +11

Aligning video generative models to human preferences heavily relies on Reinforcement Learning (RL), which suffers from extensive computational overhead. Existing workflows typical…

cs.CV2026

TextRefine: Improving Textual Fidelity, Spatial Placement, and Glyph Rendering for Text Editing in Product Posters

Honglie Wang, Jia Sun, Zijun Li +11

Text editing in product posters entails inserting new text or replacing existing text while preserving product appearance, background content, and global composition. Despite recen…

cs.CV2026

DetailAnywhere: Fashion Detail Generation via Cross-Modal Feature Alignment Distillation

Zijun Li, Yimin Zhou, Jia Sun +12

Diffusion-based generative AI has achieved remarkable success in e-commerce applications such as virtual try-on, poster generation, and product background synthesis. However, when…

cs.CV2026

CaC: Advancing Video Reward Models via Hierarchical Spatiotemporal Concentrating

Jiyuan Wang, Huan Ouyang, Jiuzhou Lin +15

In this paper, we propose Concentrate and Concentrate (CaC), a coarse-to-fine anomaly reward model based on Vision-Language Models. During inference, it first conducts a global tem…

cs.CV2026

OmniDiT: Extending Diffusion Transformer to Omni-VTON Framework

Weixuan Zeng, Pengcheng Wei, Huaiqing Wang +8

Despite the rapid advancement of Virtual Try-On (VTON) and Try-Off (VTOFF) technologies, existing VTON methods face challenges with fine-grained detail preservation, generalization…

cs.CV2024

EVLM: An Efficient Vision-Language Model for Visual Understanding

Kaibing Chen, Dong Shen, Hanwen Zhong +14

In the field of multi-modal language models, the majority of methods are built on an architecture similar to LLaVA. These models use a single-layer ViT feature as a visual prompt,…