17 papers
MVEI & EmObserver: Empowering MLLM-Oriented Visual Emotional Intelligence via Emotion Statement Judgement
Daiqing Wu, Dongbao Yang, Jiashu Yao +4
Affective Image Content Analysis (AICA) aims to recognize and understand emotions elicited by visual content, representing an indispensable step toward Artificial General Intellige…
Orthogonal Knowledge Refreshing for Domain-Incremental Object Detection
Aoting Zhang, Dongbao Yang, Chang Liu +3
Domain-incremental object detection (DIOD) requires models to continually adapt to new domains while preserving prior knowledge. Recently, parameter-efficient fine-tuning offers a…
UniTranslator: A Unified Multi-modal Framework for End-to-end In-Image Machine Translation
Jiahao Lyu, Pei Fu, Zhenhang Li +6
In-Image Machine Translation (IIMT) aims to translate scene text in an image and render the translated text back into the original regions while preserving the overall visual appea…
Chronicles-OCR: A Cross-Temporal Perception Benchmark for the Evolutionary Trajectory of Chinese Characters
Gengluo Li, Shangpin Peng, Xingyu Wan +16
Vision Large Language Models (VLLMs) have achieved remarkable success in modern text-rich visual understanding. However, their perceptual robustness in the face of the continuous m…
Learn where to Click from Yourself: On-Policy Self-Distillation for GUI Grounding
Yan Zhang, Daiqing Wu, Huawen Shen +2
Graphical User Interface (GUI) grounding maps natural language instructions to the visual coordinates of target elements and serves as a core capability for autonomous GUI agents.…
Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware Training
Gengluo Li, Pengyuan Lyu, Chengquan Zhang +7
Document parsing has recently advanced with multimodal large language models (MLLMs) that directly map document images to structured outputs. Traditional cascaded pipelines depend…