24 citations · 138 across the 68 of their papers we have counts for
72 papers · 1 filter
From Glance to Scrutiny: Progressive Distortion Reasoning for Fine-Grained Image Quality Assessment
Aoting Zhang, Mingze Gao, Dongbao Yang +5
Multi-modal large language models (MLLMs) have demonstrated significant potential in image quality assessment (IQA) by bridging visual perception with descriptive evaluations. Howe…
Knowing Beyond the Known: Reinforced Knowledge Specification for Multi-Label Class-Incremental Learning
Aoting Zhang, Dongbao Yang, Chang Liu +3
Existing class-incremental learning methods struggle in multi-label scenarios (MLCIL) due to the inherent contradiction of learning objectives arising from co-occurring and incompl…
HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better
Gengluo Li, Xingyu Wan, Shangpin Peng +20
We present HunyuanOCR-1.5, a lightweight end-to-end OCR-specialized vision-language model. HunyuanOCR unifies document parsing, text spotting, information extraction, text-image tr…
Open-Set Visual Text Forensics via Sparse-Constraint Rectified Flow
Jiangling Zhang, Shuxuan Gao, Zeyu Chen +2
Rapidly evolving Generative AI enables sophisticated visual text manipulations that increasingly evade current forensic detectors. Existing discriminative models often overfit spec…
MVEI & EmObserver: Empowering MLLM-Oriented Visual Emotional Intelligence via Emotion Statement Judgement
Daiqing Wu, Dongbao Yang, Jiashu Yao +4
Affective Image Content Analysis (AICA) aims to recognize and understand emotions elicited by visual content, representing an indispensable step toward Artificial General Intellige…
Orthogonal Knowledge Refreshing for Domain-Incremental Object Detection
Aoting Zhang, Dongbao Yang, Chang Liu +3
Domain-incremental object detection (DIOD) requires models to continually adapt to new domains while preserving prior knowledge. Recently, parameter-efficient fine-tuning offers a…