2 papers
cs.CV2025
OmniPT: Unleashing the Potential of Large Vision Language Models for Pedestrian Tracking and Understanding
Teng Fu, Mengyang Zhao, Ke Niu +2
LVLMs have been shown to perform excellently in image-level tasks such as VQA and caption. However, in many instance-level tasks, such as visual grounding and object detection, LVL…
cs.CV2025
Interpretable Oracle Bone Script Decipherment through Radical and Pictographic Analysis with LVLMs
Kaixin Peng, Mengyang Zhao, Haiyang Yu +2
As the oldest mature writing system, Oracle Bone Script (OBS) has long posed significant challenges for archaeological decipherment due to its rarity, abstractness, and pictographi…