6 papers
CA-GCL: Cross-Anatomy Global-Local Contrastive Learning for Robust 3D Medical Image Understanding
Hanwen Zhang, Yao Liu, Die Dai +4
Fine-grained Vision-Language Pre-training (FVLP) demonstrates significant potential in 3D medical image understanding by aligning anatomy-level visual representations with correspo…
UNISON: A Unified Sound Generation and Editing Framework via Deep LLM Fusion
Zhaoqing Li, Haoning Xu, Jingran Su +9
We present UNISON, a latent diffusion framework that unifies speech generation, sound generation, and audio editing within a single model. A single model handles text-to-audio, tex…
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm
Yaofang Liu, Kangning Cui, Meng Chu +7
Humans often specify and create through visual artifacts: typography sheets, sketches, reference images, and annotated scenes. Yet modern visual generators still ask users to seria…
Pusa V1.0: Unlocking Temporal Control in Pretrained Video Diffusion Models via Vectorized Timestep Adaptation
Yaofang Liu, Yumeng Ren, Aitor Artola +9
The rapid advancement of video diffusion models has been hindered by fundamental limitations in temporal modeling, particularly the rigid synchronization of frame evolution imposed…
GHPO: Adaptive Guidance for Stable and Efficient LLM Reinforcement Learning
Ziru Liu, Cheng Gong, Xinyu Fu +7
Reinforcement Learning with Verifiable Rewards (RLVR) has recently emerged as a powerful paradigm for facilitating the self-improvement of large language models (LLMs), particularl…
Improving Diffusion Generative Models via Truncated Karhunen--Loève Expansion
Yumeng Ren, Yaofang Liu, Aitor Artola +3
Pretrained diffusion models exhibit a well-known training-sampling mismatch, often attributed to exposure bias and related distribution-shift effects. We provide a quantitative int…