4 papers · 1 filter
CardioLens: Revealing the Clinical Reality Gap of MLLMs via Multi-Sequence Cardiac MRI Evaluations
Zixian Su, Hongkai Zhang, Fan Gao +12
Multimodal Large Language Models (MLLMs) have shown strong performance on public medical benchmarks, yet existing evaluations often remain weak proxies for clinical use, relying on…
Beyond Ground-Truth: Leveraging Image Quality Priors for Real-World Image Restoration
Fengyang Xiao, Peng Hu, Lei Xu +7
Real-world image restoration aims to restore high-quality (HQ) images from degraded low-quality (LQ) inputs captured under uncontrolled conditions. Existing methods typically depen…
Text to Sketch Generation with Multi-Styles
Tengjie Li, Shikui Tu, Lei Xu
Recent advances in vision-language models have facilitated progress in sketch generation. However, existing specialized methods primarily focus on generic synthesis and lack mechan…
PFB-Diff: Progressive Feature Blending Diffusion for Text-driven Image Editing
Wenjing Huang, Shikui Tu, Lei Xu
Diffusion models have demonstrated their ability to generate diverse and high-quality images, sparking considerable interest in their potential for real image editing applications.…