5 papers
Copula-enhanced Vision Transformer for high myopia diagnosis through OU UWF fundus images
Chong Zhong, Yunhao Liu, Yang Li +10
The advancement of AI-assisted myopia screening necessitates the joint diagnosis of both-eye (OU) high myopia (HM) status and the prediction of axial length (AL). This clinical req…
Rethinking Soft Compression in Retrieval-Augmented Generation: A Query-Conditioned Selector Perspective
Yunhao Liu, Zian Jia, Xinyu Gao +2
Retrieval-Augmented Generation (RAG) effectively grounds Large Language Models (LLMs) with external knowledge and is widely applied to Web-related tasks. However, its scalability i…
Reasoning with Autoregressive-Diffusion Collaborative Thoughts
Mu Yuan, Liekang Zeng, Guoliang Xing +2
Autoregressive and diffusion models represent two complementary generative paradigms. Autoregressive models excel at sequential planning and constraint composition, yet struggle wi…
Taming a Retrieval Framework to Read Images in Humanlike Manner for Augmenting Generation of MLLMs
Suyang Xi, Chenxi Yang, Hong Ding +4
Multimodal large language models (MLLMs) often fail in fine-grained visual question answering, producing hallucinations about object identities, positions, and relations because te…
Multimodal Medical Image Binding via Shared Text Embeddings
Yunhao Liu, Suyang Xi, Shiqi Liu +6
Medical image analysis increasingly relies on the integration of multiple imaging modalities to capture complementary anatomical and functional information, enabling more accurate…