5 papers
Avatar V: Scaling Video-Reference Avatar Video Generation
Benjamin Liang, Ce Chen, Desmond Lin +20
Generating avatar videos that are not merely visually similar to a target individual but behaviorally recognizable, faithfully reproducing their talking rhythm, gestural tendencies…
TransVLM: A Vision-Language Framework and Benchmark for Detecting Any Shot Transitions
Ce Chen, Yi Ren, Yuanming Li +5
Traditional Shot Boundary Detection (SBD) inherently struggles with complex transitions by formulating the task around isolated cut points, frequently yielding corrupted video shot…
Generate Your Talking Avatar from Video Reference
Zujin Guo, Zhenhui Ye, Yi Ren +4
Existing talking avatar methods typically adopt an image-to-video pipeline conditioned on a static reference image within the same scene as the target generation. This restricted,…
Mobile-VTON: High-Fidelity On-Device Virtual Try-On
Zhenchen Wan, Ce Chen, Runqi Lin +5
Virtual try-on (VTON) has recently achieved impressive visual fidelity, but most existing systems require uploading personal photos to cloud-based GPUs, raising privacy concerns an…
Condition Matters in Full-head 3D GANs
Heyuan Li, Huimin Zhang, Yuda Qiu +10
Conditioning is crucial for stable training of full-head 3D GANs. Without any conditioning signal, the model suffers from severe mode collapse, making it impractical to training. H…