papers

Publications (75)

cs.CL2026

MR-Align: Meta-Reasoning Informed Factuality Alignment for Large Reasoning Models

Xinming Wang, Jian Xu, Bin Yu +9

Large reasoning models (LRMs) show strong capabilities in complex reasoning, yet their marginal gains on evidence-dependent factual questions are limited. We find this limitation i…

cs.LG2023

Dynamics-Aware Loss for Learning with Label Noise

Xiu-Chuan Li, Xiaobo Xia, Fei Zhu +3

Label noise poses a serious threat to deep neural networks (DNNs). Employing robust loss functions which reconcile fitting ability with robustness is a simple but effective strateg…

cs.AI2025

MV-MATH: Evaluating Multimodal Math Reasoning in Multi-Visual Contexts

Peijie Wang, Zhong-Zhi Li, Fei Yin +3

Multimodal Large Language Models (MLLMs) have shown promising capabilities in mathematical reasoning within visual contexts across various datasets. However, most existing multimod…

cs.CV2024

PASS++: A Dual Bias Reduction Framework for Non-Exemplar Class-Incremental Learning

Fei Zhu, Xu-Yao Zhang, Zhen Cheng +1

Class-incremental learning (CIL) aims to recognize new classes incrementally while maintaining the discriminability of old classes. Most existing CIL methods are exemplar-based, i.…

cs.AI2026

SymDiag: Explainable Diagnosis for LLM Reasoning via Neuro-Symbolic Verification

Wenyao Cui, Huaping Zhang, Yongyi Huang +6

Large language models (LLMs) increasingly serve as data-driven reasoners, yet their chains-of-thought (CoT) can be unfaithful even when final answers are correct. Most existing ``v…

cs.AI2025

From System 1 to System 2: A Survey of Reasoning Large Language Models

Zhong-Zhi Li, Duzhen Zhang, Ming-Liang Zhang +18

Achieving human-level intelligence requires refining the transition from the fast, intuitive System 1 to the slower, more deliberate System 2 reasoning. While System 1 excels in qu…

cs.CV2023

Class Incremental Learning with Self-Supervised Pre-Training and Prototype Learning

Wenzhuo Liu, Xinjian Wu, Fei Zhu +3

Deep Neural Network (DNN) has achieved great success on datasets of closed class set. However, new classes, like new categories of social media topics, are continuously added to th…

cs.LG2023

Rethinking Confidence Calibration for Failure Prediction

Fei Zhu, Zhen Cheng, Xu-Yao Zhang +1

Reliable confidence estimation for the predictions is important in many safety-critical applications. However, modern deep neural networks are often overconfident for their incorre…

cs.CV2023

ICDAR 2023 Competition on Structured Text Extraction from Visually-Rich Document Images

Wenwen Yu, Chengquan Zhang, Haoyu Cao +24

Structured text extraction is one of the most valuable and challenging application directions in the field of Document AI. However, the scenarios of past benchmarks are limited, an…

cs.CV2024

Differentiable Proximal Graph Matching

Haoru Tan, Chuang Wang, Xu-Yao Zhang +1

Graph matching is a fundamental tool in computer vision and pattern recognition. In this paper, we introduce an algorithm for graph matching based on the proximal operator, referre…

cs.LG2023

OpenMix: Exploring Outlier Samples for Misclassification Detection

Fei Zhu, Zhen Cheng, Xu-Yao Zhang +1

Reliable confidence estimation for deep neural classifiers is a challenging yet fundamental requirement in high-stakes applications. Unfortunately, modern deep neural networks are…

cs.CV2024

Unified Entropy Optimization for Open-Set Test-Time Adaptation

Zhengqing Gao, Xu-Yao Zhang, Cheng-Lin Liu

Test-time adaptation (TTA) aims at adapting a model pre-trained on the labeled source domain to the unlabeled target domain. Existing methods usually focus on improving TTA perform…

cs.CV2016

Natural Scene Character Recognition Using Robust PCA and Sparse Representation

Zheng Zhang, Yong Xu, Cheng-Lin Liu

Natural scene character recognition is challenging due to the cluttered background, which is hard to separate from text. In this paper, we propose a novel method for robust scene c…

cs.AI2026

BiNSGPS: Geometry Problem Solving via Bidirectional Neuro-Symbolic Interaction

Qi Wang, Peijie Wang, Fei Yin +1

Geometry problem solving poses distinct challenges in artificial intelligence. Existing approaches typically fall into two paradigms: symbolic methods, which exhibit limited adapta…

cs.CV2016

Drawing and Recognizing Chinese Characters with Recurrent Neural Network

Xu-Yao Zhang, Fei Yin, Yan-Ming Zhang +2

Recent deep learning based approaches have achieved great success on handwriting recognition. Chinese characters are among the most widely adopted writing systems in the world. Pre…

cs.CV2016

Online and Offline Handwritten Chinese Character Recognition: A Comprehensive Study and New Benchmark

Xu-Yao Zhang, Yoshua Bengio, Cheng-Lin Liu

Recent deep learning based methods have achieved the state-of-the-art performance for handwritten Chinese character recognition (HCCR) by learning discriminative representations di…

cs.LG2025

ProtoGCD: Unified and Unbiased Prototype Learning for Generalized Category Discovery

Shijie Ma, Fei Zhu, Xu-Yao Zhang +1

Generalized category discovery (GCD) is a pragmatic but underexplored problem, which requires models to automatically cluster and discover novel categories by leveraging the labele…

cs.LG2025

Federated Continual Instruction Tuning

Haiyang Guo, Fanhu Zeng, Fei Zhu +5

A vast amount of instruction tuning data is crucial for the impressive performance of Large Multimodal Models (LMMs), but the associated computational costs and data collection dem…

cs.LG2018

Anomaly Detection via Minimum Likelihood Generative Adversarial Networks

Chu Wang, Yan-Ming Zhang, Cheng-Lin Liu

Anomaly detection aims to detect abnormal events by a model of normality. It plays an important role in many domains such as network intrusion detection, criminal activity identity…

cs.LG2025

Open-world machine learning: A review and new outlooks

Fei Zhu, Shijie Ma, Zhen Cheng +4

Machine learning has achieved remarkable success in many applications. However, existing studies are largely based on the closed-world assumption, which assumes that the environmen…

cs.CV2026

A Lightweight Context-Driven Training-Free Network for Scene Text Segmentation and Recognition

Ritabrata Chakraborty, Shivakumara Palaiahnakote, Umapada Pal +1

Modern scene text recognition systems often depend on large end-to-end architectures that require extensive training and are prohibitively expensive for real-time scenarios. In suc…

cs.CV2026

Geoparsing: Diagram Parsing for Plane and Solid Geometry with a Unified Formal Language

Peijie Wang, Ming-Liang Zhang, Jun Cao +10

Multimodal Large Language Models (MLLMs) have achieved remarkable progress but continue to struggle with geometric reasoning, primarily due to the perception bottleneck regarding f…

cs.CV2024

Towards Non-Exemplar Semi-Supervised Class-Incremental Learning

Wenzhuo Liu, Fei Zhu, Cheng-Lin Liu

Deep neural networks perform remarkably well in close-world scenarios. However, novel classes emerged continually in real applications, making it necessary to learn incrementally.…

cs.CV2022

Towards Open-Set Text Recognition via Label-to-Prototype Learning

Chang Liu, Chun Yang, Hai-Bo Qin +3

Scene text recognition is a popular topic and extensively used in the industry. Although many methods have achieved satisfactory performance for the close-set text recognition chal…

cs.CV2024

Multi-scale Unified Network for Image Classification

Wenzhuo Liu, Fei Zhu, Cheng-Lin Liu

Convolutional Neural Networks (CNNs) have advanced significantly in visual representation learning and recognition. However, they face notable challenges in performance and computa…

cs.CV2018

Robust Classification with Convolutional Prototype Learning

Hong-Ming Yang, Xu-Yao Zhang, Fei Yin +1

Convolutional neural networks (CNNs) have been widely used for image classification. Despite its high accuracies, CNN has been shown to be easily fooled by some adversarial example…

cs.AI2026

HVR-Met: A Hypothesis-Verification-Replanning Agentic System for Extreme Weather Diagnosis

Shuo Tang, Jiadong Zhang, Gengxian Zhou +11

While deep learning-based weather forecasting paradigms have made significant strides, addressing extreme weather diagnostics remains a formidable challenge. This gap exists primar…

cs.CV2024

Active Generalized Category Discovery

Shijie Ma, Fei Zhu, Zhun Zhong +2

Generalized Category Discovery (GCD) is a pragmatic and challenging open-world task, which endeavors to cluster unlabeled samples from both novel and old classes, leveraging some l…

cs.CV2017

Scene Text Recognition with Sliding Convolutional Character Models

Fei Yin, Yi-Chao Wu, Xu-Yao Zhang +1

Scene text recognition has attracted great interests from the computer vision and pattern recognition community in recent years. State-of-the-art methods use concolutional neural n…

cs.CL2025

HiDe-LLaVA: Hierarchical Decoupling for Continual Instruction Tuning of Multimodal Large Language Model

Haiyang Guo, Fanhu Zeng, Ziwei Xiang +4

Instruction tuning is widely used to improve a pre-trained Multimodal Large Language Model (MLLM) by training it on curated task-specific datasets, enabling better comprehension of…

cs.CV2024

MSPE: Multi-Scale Patch Embedding Prompts Vision Transformers to Any Resolution

Wenzhuo Liu, Fei Zhu, Shijie Ma +1

Although Vision Transformers (ViTs) have recently advanced computer vision tasks significantly, an important real-world problem was overlooked: adapting to variable input resolutio…

cs.CV2022

Document Dewarping with Control Points

Guo-Wang Xie, Fei Yin, Xu-Yao Zhang +1

Document images are now widely captured by handheld devices such as mobile phones. The OCR performance on these images are largely affected due to geometric distortion of the docum…

cs.CV2012

A Fast Projected Fixed-Point Algorithm for Large Graph Matching

Yao Lu, Kaizhu Huang, Cheng-Lin Liu

We propose a fast approximate algorithm for large graph matching. A new projected fixed-point method is defined and a new doubly stochastic projection is adopted to derive the algo…

cs.CV2024

Recoverable Compression: A Multimodal Vision Token Recovery Mechanism Guided by Text Information

Yi Chen, Jian Xu, Xu-Yao Zhang +3

With the advancement of large-scale language modeling techniques, large multimodal models combining visual encoders with large language models have demonstrated exceptional perform…

cs.AI2025

Multimodal Agricultural Agent Architecture (MA3): A New Paradigm for Intelligent Agricultural Decision-Making

Zhuoning Xu, Jian Xu, Mingqing Zhang +3

As a strategic pillar industry for human survival and development, modern agriculture faces dual challenges: optimizing production efficiency and achieving sustainable development.…

cs.CV2024

Unified Classification and Rejection: A One-versus-All Framework

Zhen Cheng, Xu-Yao Zhang, Cheng-Lin Liu

Classifying patterns of known classes and rejecting ambiguous and novel (also called as out-of-distribution (OOD)) inputs are involved in open world pattern recognition. Deep neura…

cs.CV2024

Happy: A Debiased Learning Framework for Continual Generalized Category Discovery

Shijie Ma, Fei Zhu, Zhun Zhong +3

Constantly discovering novel concepts is crucial in evolving environments. This paper explores the underexplored task of Continual Generalized Category Discovery (C-GCD), which aim…

cs.AI2025

LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and Locating

Chao Deng, Jiale Yuan, Pi Bu +8

Large vision language models (LVLMs) have improved the document understanding capabilities remarkably, enabling the handling of complex document elements, longer contexts, and a wi…

cs.NE2024

Biologically Plausible Training of Deep Neural Networks Using a Top-down Credit Assignment Network

Jian-Hui Chen, Cheng-Lin Liu, Zuoren Wang

Despite the widespread adoption of Backpropagation algorithm-based Deep Neural Networks, the biological infeasibility of the BP algorithm could potentially limit the evolution of n…

cs.AI2026

CatalyticMLLM: A Graph-Text Multimodal Large Language Model for Catalytic Materials

Yanjie Li, Jian Xu, Xu-Yao Zhang +4

Property prediction and inverse structural design of catalytic materials are typically modeled as two independent tasks: the former predicts target properties from given structures…

cs.CV2017

Deep Direct Regression for Multi-Oriented Scene Text Detection

Wenhao He, Xu-Yao Zhang, Fei Yin +1

In this paper, we first provide a new perspective to divide existing high performance object detection methods into direct and indirect regressions. Direct regression performs boun…

cs.CV2021

Dewarping Document Image By Displacement Flow Estimation with Fully Convolutional Network

Guo-Wang Xie, Fei Yin, Xu-Yao Zhang +1

As camera-based documents are increasingly used, the rectification of distorted document images becomes a need to improve the recognition performance. In this paper, we propose a n…

cs.CV2024

PILoRA: Prototype Guided Incremental LoRA for Federated Class-Incremental Learning

Haiyang Guo, Fei Zhu, Wenzhuo Liu +2

Existing federated learning methods have effectively dealt with decentralized learning in scenarios involving data privacy and non-IID data. However, in real-world situations, each…

cs.CV2018

Practical Block-wise Neural Network Architecture Generation

Zhao Zhong, Junjie Yan, Wei Wu +2

Convolutional neural networks have gained a remarkable success in computer vision. However, most usable network architectures are hand-crafted and usually require expertise and ela…

cs.CV2024

Revisiting Confidence Estimation: Towards Reliable Failure Prediction

Fei Zhu, Xu-Yao Zhang, Zhen Cheng +1

Reliable confidence estimation is a challenging yet fundamental requirement in many risk-sensitive applications. However, modern deep neural networks are often overconfident for th…

cs.CV2016

Hybrid CNN and Dictionary-Based Models for Scene Recognition and Domain Adaptation

Guo-Sen Xie, Xu-Yao Zhang, Shuicheng Yan +1

Convolutional neural network (CNN) has achieved state-of-the-art performance in many different visual tasks. Learned from a large-scale training dataset, CNN features are much more…

cs.CG2025

SOLIDGEO: Measuring Multimodal Spatial Math Reasoning in Solid Geometry

Peijie Wang, Chao Yang, Zhong-Zhi Li +6

Geometry is a fundamental branch of mathematics and plays a crucial role in evaluating the reasoning capabilities of multimodal large language models (MLLMs). However, existing mul…

cs.CV2022

Unsupervised Structure-Texture Separation Network for Oracle Character Recognition

Mei Wang, Weihong Deng, Cheng-Lin Liu

Oracle bone script is the earliest-known Chinese writing system of the Shang dynasty and is precious to archeology and philology. However, real-world scanned oracle data are rare a…

cs.LG2025

Continual Learning for Generative AI: From LLMs to MLLMs and Beyond

Haiyang Guo, Fanhu Zeng, Fei Zhu +9

The rapid advancement of generative models has empowered modern AI systems to comprehend and produce highly sophisticated content, even achieving human-level performance in specifi…

cs.LG2023

Average of Pruning: Improving Performance and Stability of Out-of-Distribution Detection

Zhen Cheng, Fei Zhu, Xu-Yao Zhang +1

Detecting Out-of-distribution (OOD) inputs have been a critical issue for neural networks in the open world. However, the unstable behavior of OOD detection along the optimization…

cs.AI2024

Fuse, Reason and Verify: Geometry Problem Solving with Parsed Clauses from Diagram

Ming-Liang Zhang, Zhong-Zhi Li, Fei Yin +2

Geometry problem solving (GPS) requires capacities of multi-modal understanding, multi-hop reasoning and theorem knowledge application. In this paper, we propose a neural-symbolic…

cs.CV2022

Plane Geometry Diagram Parsing

Ming-Liang Zhang, Fei Yin, Yi-Han Hao +1

Geometry diagram parsing plays a key role in geometry problem solving, wherein the primitive extraction and relation parsing remain challenging due to the complex layout and betwee…

cs.CV2024

LANS: A Layout-Aware Neural Solver for Plane Geometry Problem

Zhong-Zhi Li, Ming-Liang Zhang, Fei Yin +1

Geometry problem solving (GPS) is a challenging mathematical reasoning task requiring multi-modal understanding, fusion, and reasoning. Existing neural solvers take GPS as a vision…

cs.CL2025

TL;DR: Too Long, Do Re-weighting for Efficient LLM Reasoning Compression

Zhong-Zhi Li, Xiao Liang, Zihao Tang +11

Large Language Models (LLMs) have recently achieved remarkable progress by leveraging Reinforcement Learning and extended Chain-of-Thought (CoT) techniques. However, the challenge…

cs.CV2022

Emergence of Machine Language: Towards Symbolic Intelligence with Neural Networks

Yuqi Wang, Xu-Yao Zhang, Cheng-Lin Liu +1

Representation is a core issue in artificial intelligence. Humans use discrete language to communicate and learn from each other, while machines use continuous features (like vecto…

cs.CV2018

BlockQNN: Efficient Block-wise Neural Network Architecture Generation

Zhao Zhong, Zichen Yang, Boyang Deng +4

Convolutional neural networks have gained a remarkable success in computer vision. However, most usable network architectures are hand-crafted and usually require expertise and ela…

cs.CV2018

SCAN: Sliding Convolutional Attention Network for Scene Text Recognition

Yi-Chao Wu, Fei Yin, Xu-Yao Zhang +2

Scene text recognition has drawn great attentions in the community of computer vision and artificial intelligence due to its challenges and wide applications. State-of-the-art recu…

cs.LG2012

Robust Metric Learning by Smooth Optimization

Kaizhu Huang, Rong Jin, Zenglin Xu +1

Most existing distance metric learning methods assume perfect side information that is usually given in pairwise or triplet constraints. Instead, in many real-world applications, t…

cs.AI2025

PDE-Agent: A toolchain-augmented multi-agent framework for PDE solving

Jianming Liu, Ren Zhu, Jian Xu +4

Solving Partial Differential Equations (PDEs) is a cornerstone of engineering and scientific research. Traditional methods for PDE solving are cumbersome, relying on manual setup a…

cs.CV2024

Adapting Vision-Language Models to Open Classes via Test-Time Prompt Tuning

Zhengqing Gao, Xiang Ao, Xu-Yao Zhang +1

Adapting pre-trained models to open classes is a challenging problem in machine learning. Vision-language models fully explore the knowledge of text modality, demonstrating strong…

cs.CV2025

Cross-Modal Causal Intervention for Medical Report Generation

Weixing Chen, Yang Liu, Ce Wang +4

Radiology Report Generation (RRG) is essential for computer-aided diagnosis and medication guidance, which can relieve the heavy burden of radiologists by automatically generating…

cs.LG2024

Branch-Tuning: Balancing Stability and Plasticity for Continual Self-Supervised Learning

Wenzhuo Liu, Fei Zhu, Cheng-Lin Liu

Self-supervised learning (SSL) has emerged as an effective paradigm for deriving general representations from vast amounts of unlabeled data. However, as real-world applications co…

cs.CV2025

LLaVA-c: Continual Improved Visual Instruction Tuning

Wenzhuo Liu, Fei Zhu, Haiyang Guo +2

Multimodal models like LLaVA-1.5 achieve state-of-the-art visual understanding through visual instruction tuning on multitask datasets, enabling strong instruction-following and mu…

cs.CV2019

Arbitrary Shape Scene Text Detection with Adaptive Text Region Representation

Xiaobing Wang, Yingying Jiang, Zhenbo Luo +3

Scene text detection attracts much attention in computer vision, because it can be widely used in many applications such as real-time text translation, automatic information entry,…

cs.CV2025

DocSAM: Unified Document Image Segmentation via Query Decomposition and Heterogeneous Mixed Learning

Xiao-Hui Li, Fei Yin, Cheng-Lin Liu

Document image segmentation is crucial for document analysis and recognition but remains challenging due to the diversity of document formats and segmentation tasks. Existing metho…

cs.CV2024

StylePrompter: Enhancing Domain Generalization with Test-Time Style Priors

Jiao Zhang, Jian Xu, Xu-Yao Zhang +1

In real-world applications, the sample distribution at the inference stage often differs from the one at the training stage, causing performance degradation of trained deep models.…

cs.CV2024

Ensemble Quadratic Assignment Network for Graph Matching

Haoru Tan, Chuang Wang, Sitong Wu +3

Graph matching is a commonly used technique in computer vision and pattern recognition. Recent data-driven approaches have improved the graph matching accuracy remarkably, whereas…

cs.CV2024

WPS-SAM: Towards Weakly-Supervised Part Segmentation with Foundation Models

Xinjian Wu, Ruisong Zhang, Jie Qin +2

Segmenting and recognizing diverse object parts is crucial in computer vision and robotics. Despite significant progress in object segmentation, part-level segmentation remains und…

cs.CV2025

ModalPrompt: Towards Efficient Multimodal Continual Instruction Tuning with Dual-Modality Guided Prompt

Fanhu Zeng, Fei Zhu, Haiyang Guo +2

Large Multimodal Models (LMMs) exhibit remarkable multi-tasking ability by learning mixed instruction datasets. However, novel tasks would be encountered sequentially in dynamic wo…

cs.CV2020

Weakly-Supervised Arbitrary-Shaped Text Detection with Expectation-Maximization Algorithm

Mengbiao Zhao, Wei Feng, Fei Yin +2

Arbitrary-shaped text detection is an important and challenging task in computer vision. Most existing methods require heavy data labeling efforts to produce polygon-level text reg…

cs.CL2024

CMMaTH: A Chinese Multi-modal Math Skill Evaluation Benchmark for Foundation Models

Zhong-Zhi Li, Ming-Liang Zhang, Fei Yin +7

Due to the rapid advancements in multimodal large language models, evaluating their multimodal mathematical capabilities continues to receive wide attention. Despite the datasets l…

cs.CV2020

Towards Robust Pattern Recognition: A Review

Xu-Yao Zhang, Cheng-Lin Liu, Ching Y. Suen

The accuracies for many pattern recognition tasks have increased rapidly year by year, achieving or even outperforming human performance. From the perspective of accuracy, pattern…

cs.AI2023

A Multi-Modal Neural Geometric Solver with Textual Clauses Parsed from Diagram

Ming-Liang Zhang, Fei Yin, Cheng-Lin Liu

Geometry problem solving (GPS) is a high-level mathematical reasoning requiring the capacities of multi-modal fusion and geometric knowledge application. Recently, neural solvers h…

cs.LG2025

Air Quality Prediction with A Meteorology-Guided Modality-Decoupled Spatio-Temporal Network

Hang Yin, Yan-Ming Zhang, Jian Xu +3

Air quality prediction plays a crucial role in public health and environmental protection. Accurate air quality prediction is a complex multivariate spatiotemporal problem, that in…

cs.CV2023

Towards Reliable Domain Generalization: A New Dataset and Evaluations

Jiao Zhang, Xu-Yao Zhang, Cheng-Lin Liu

There are ubiquitous distribution shifts in the real world. However, deep neural networks (DNNs) are easily biased towards the training set, which causes severe performance degrada…