2 papers
cs.CV2026
Entity-Faithful Repair of Synthetic Supervision for Zero-Shot Image Captioning
Zhiyue Liu, Wenkai Zhou, Jian Qin +1
Zero-shot image captioning aims to generate image descriptions without annotated image-text pairs. Recent approaches exploit text-to-image models to synthesize training data from t…
cs.CV2025
OralGPT-Omni: A Versatile Dental Multimodal Large Language Model
Jing Hao, Yuci Liang, Lizhuo Lin +12
Multimodal Large Language Models (MLLMs) have exhibited immense potential across numerous medical specialties; yet, dentistry remains underexplored, in part due to limited domain-s…