4 papers
Energy-oriented Diffusion Bridge for Image Restoration with Foundational Diffusion Models
Jinhui Hou, Zhiyu Zhu, Junhui Hou
Diffusion bridge models have shown great promise in image restoration by explicitly connecting clean and degraded image distributions. However, they often rely on complex and high-…
Cross-Modal Coreference Alignment: Enabling Reliable Information Transfer in Omni-LLMs
Hongcheng Liu, Yuhao Wang, Zhe Chen +5
Omni Large Language Models (Omni-LLMs) have demonstrated impressive capabilities in holistic multi-modal perception, yet they consistently falter in complex scenarios requiring syn…
HeteroRAG: A Heterogeneous Retrieval-Augmented Generation Framework for Medical Vision Language Tasks
Zhe Chen, Yusheng Liao, Zhiyuan Zhu +4
Medical large vision-language Models (Med-LVLMs) have shown promise in clinical applications but suffer from factual inaccuracies and unreliable outputs, posing risks in real-world…
Drawing the Line: Enhancing Trustworthiness of MLLMs Through the Power of Refusal
Yuhao Wang, Zhiyuan Zhu, Heyang Liu +4
Multimodal large language models (MLLMs) excel at multimodal perception and understanding, yet their tendency to generate hallucinated or inaccurate responses undermines their trus…