activity
20242026
collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2026

RAGDiffusion++: From Macro-Retrieval to Micro-Fidelity Alignment for Garment Generation

Yuhan Li, Xianfeng Tan, Fangao Zeng +6

Standard clothing asset generation---restoring forward-facing flat-lay garment images from diverse real-world contexts---holds immense commercial value yet demands both macroscopic…

cs.CV2026

RL-Native Distillation: Exploiting Scored Trajectories for Few-Step Image Generation

Yuhan Li, Fangao Zeng, Sicong Kang +5

Efficient text-to-image generation requires both reinforcement-learning (RL)-based reward alignment and few-step distillation, yet these procedures are typically performed sequenti…

cs.CV2026

TRUST: Efficient Abdominal Trauma Recognition via Image-to-Ultrasound-Video Transfer Learning

Enguang Wang, Hao Zhou, Shuo Gao +2

Abdominal ultrasound is indispensable for rapid, noninvasive trauma triage. However, interpreting the subtle dynamic cues embedded in continuous scanning is time-intensive and oper…

cs.CV2024

Video Creation by Demonstration

Yihong Sun, Hao Zhou, Liangzhe Yuan +7

We explore a novel video creation experience, namely Video Creation by Demonstration. Given a demonstration video and a context image from a different scene, we generate a physical…

cs.CV2024

Open-Vocabulary 3D Semantic Segmentation with Text-to-Image Diffusion Models

Xiaoyu Zhu, Hao Zhou, Pengfei Xing +6

In this paper, we investigate the use of diffusion models which are pre-trained on large-scale image-caption pairs for open-vocabulary 3D semantic understanding. We propose a novel…