2 papers
cs.CV2026
Latent Ordinal Evidence, Misaligned Outputs: Inference-Time Ordinal Lens Alignment for Multimodal LLMs
Haiming Li, Yingsheng Liu, Jingmin Zhu +5
Multimodal LLMs apply the language model interface to visual inputs, where ordinal regression tasks such as age estimation, image quality assessment, and disease grading require au…
cs.CV2026
Unlocking the Power of Medical Tabular Data via Semantic-Aware Multimodal Pre-training
Yingsheng Liu, Haiming Li, Jingmin Zhu +6
While vision-language models dominate medical representation learning, unstructured text lacks the dense, quantitative diagnostic phenotypes inherent in structured clinical tables.…