6 papers
Beyond Single Character: Evaluating MLLMs for Sentence-Level Oracle Bone Inscription Understanding
Ziqi Li, Zijian Chen, Tingzhu Chen +1
Existing AI-assisted oracle bone inscription (OBI) visual recognition and understanding studies mainly focus on character-level, ignoring the long-form textual coherence and contex…
GTPred: Benchmarking MLLMs for Interpretable Geo-localization and Time-of-capture Prediction
Jinnao Li, Zijian Chen, Tingzhu Chen +1
Geo-localization aims to infer the geographic location where an image was captured using observable visual evidence. Traditional methods achieve impressive results through large-sc…
Find Them All: Unveiling MLLMs for Versatile Person Re-identification
Jinhao Li, Zijian Chen, Lirong Deng +2
Person re-identification (ReID) aims to retrieve images of a target person from the gallery set, with wide applications in medical rehabilitation and public security. However, trad…
Hulu-Med: A Transparent Generalist Model towards Holistic Medical Vision-Language Understanding
Songtao Jiang, Yuan Wang, Sibo Song +22
Real-world clinical decision-making requires integrating heterogeneous data, including medical text, 2D images, 3D volumes, and videos, while existing AI systems fail to unify all…
PictOBI-20k: Unveiling Large Multimodal Models in Visual Decipherment for Pictographic Oracle Bone Characters
Zijian Chen, Wenjie Hua, Jinhao Li +4
Deciphering oracle bone characters (OBCs), the oldest attested form of written Chinese, has remained the ultimate, unwavering goal of scholars, offering an irreplaceable key to und…
OBIFormer: A Fast Attentive Denoising Framework for Oracle Bone Inscriptions
Jinhao Li, Zijian Chen, Tingzhu Chen +2
Oracle bone inscriptions (OBIs) are the earliest known form of Chinese characters and serve as a valuable resource for research in anthropology and archaeology. However, most excav…