4 papers
POLARIS: Projection-Orthogonal Least Squares for Robust and Adaptive Inversion in Diffusion Models
Wenshuo Chen, Haosen Li, Shaofeng Liang +6
The Inversion-Denoising Paradigm, which is based on diffusion models, excels in diverse image editing and restoration tasks. We revisit its mechanism and reveal a critical, overloo…
CoEmoGen: Towards Semantically-Coherent and Scalable Emotional Image Content Generation
Kaishen Yuan, Yuting Zhang, Shang Gao +3
Emotional Image Content Generation (EICG) aims to generate semantically clear and emotionally faithful images based on given emotion categories, with broad application prospects. W…
MedTVT-R1: A Multimodal LLM Empowering Medical Reasoning and Diagnosis
Yuting Zhang, Kaishen Yuan, Hao Lu +3
Accurate and interpretable multi-disease diagnosis remains a critical challenge in medical research, particularly when leveraging heterogeneous multimodal medical data. Current app…
ANT: Adaptive Neural Temporal-Aware Text-to-Motion Model
Wenshuo Chen, Kuimou Yu, Haozhe Jia +8
While diffusion models advance text-to-motion generation, their static semantic conditioning ignores temporal-frequency demands: early denoising requires structural semantics for m…