4 papers
LD-LAudio-V1: Video-to-Long-Form-Audio Generation Extension with Dual Lightweight Adapters
Haomin Zhang, Kristin Qi, Shuxin Yang +3
Generating high-quality and temporally synchronized audio from video content is essential for video editing and post-production tasks, enabling the creation of semantically aligned…
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1
Xinhan Di, Kristin Qi, Pengqian Yu
Recent advances in diffusion-based video generation have enabled photo-realistic short clips, but current methods still struggle to achieve multi-modal consistency when jointly gen…
Attentional Triple-Encoder Network in Spatiospectral Domains for Medical Image Segmentation
Kristin Qi, Xinhan Di
Retinal Optical Coherence Tomography (OCT) segmentation is essential for diagnosing pathology. Traditional methods focus on either spatial or spectral domains, overlooking their co…
Towards Full-parameter and Parameter-efficient Self-learning For Endoscopic Camera Depth Estimation
Shuting Zhao, Chenkang Du, Kristin Qi +2
Adaptation methods are developed to adapt depth foundation models to endoscopic depth estimation recently. However, such approaches typically under-perform training since they limi…