3 papers
cs.CV2026
Latent Ordinal Evidence, Misaligned Outputs: Inference-Time Ordinal Lens Alignment for Multimodal LLMs
Haiming Li, Yingsheng Liu, Jingmin Zhu +5
Multimodal LLMs apply the language model interface to visual inputs, where ordinal regression tasks such as age estimation, image quality assessment, and disease grading require au…
cs.CV2026
Unlocking the Power of Medical Tabular Data via Semantic-Aware Multimodal Pre-training
Yingsheng Liu, Haiming Li, Jingmin Zhu +6
While vision-language models dominate medical representation learning, unstructured text lacks the dense, quantitative diagnostic phenotypes inherent in structured clinical tables.…
cs.CV2025
Controllable Skin Synthesis via Lesion-Focused Vector Autoregression Model
Jiajun Sun, Zhen Yu, Siyuan Yan +3
Skin images from real-world clinical practice are often limited, resulting in a shortage of training data for deep-learning models. While many studies have explored skin image synt…