Publications (75)
MR-Align: Meta-Reasoning Informed Factuality Alignment for Large Reasoning Models
Xinming Wang, Jian Xu, Bin Yu +9
Large reasoning models (LRMs) show strong capabilities in complex reasoning, yet their marginal gains on evidence-dependent factual questions are limited. We find this limitation i…
Dynamics-Aware Loss for Learning with Label Noise
Xiu-Chuan Li, Xiaobo Xia, Fei Zhu +3
Label noise poses a serious threat to deep neural networks (DNNs). Employing robust loss functions which reconcile fitting ability with robustness is a simple but effective strateg…
MV-MATH: Evaluating Multimodal Math Reasoning in Multi-Visual Contexts
Peijie Wang, Zhong-Zhi Li, Fei Yin +3
Multimodal Large Language Models (MLLMs) have shown promising capabilities in mathematical reasoning within visual contexts across various datasets. However, most existing multimod…
PASS++: A Dual Bias Reduction Framework for Non-Exemplar Class-Incremental Learning
Fei Zhu, Xu-Yao Zhang, Zhen Cheng +1
Class-incremental learning (CIL) aims to recognize new classes incrementally while maintaining the discriminability of old classes. Most existing CIL methods are exemplar-based, i.…
SymDiag: Explainable Diagnosis for LLM Reasoning via Neuro-Symbolic Verification
Wenyao Cui, Huaping Zhang, Yongyi Huang +6
Large language models (LLMs) increasingly serve as data-driven reasoners, yet their chains-of-thought (CoT) can be unfaithful even when final answers are correct. Most existing ``v…
From System 1 to System 2: A Survey of Reasoning Large Language Models
Zhong-Zhi Li, Duzhen Zhang, Ming-Liang Zhang +18
Achieving human-level intelligence requires refining the transition from the fast, intuitive System 1 to the slower, more deliberate System 2 reasoning. While System 1 excels in qu…
Class Incremental Learning with Self-Supervised Pre-Training and Prototype Learning
Wenzhuo Liu, Xinjian Wu, Fei Zhu +3
Deep Neural Network (DNN) has achieved great success on datasets of closed class set. However, new classes, like new categories of social media topics, are continuously added to th…
Rethinking Confidence Calibration for Failure Prediction
Fei Zhu, Zhen Cheng, Xu-Yao Zhang +1
Reliable confidence estimation for the predictions is important in many safety-critical applications. However, modern deep neural networks are often overconfident for their incorre…
ICDAR 2023 Competition on Structured Text Extraction from Visually-Rich Document Images
Wenwen Yu, Chengquan Zhang, Haoyu Cao +24
Structured text extraction is one of the most valuable and challenging application directions in the field of Document AI. However, the scenarios of past benchmarks are limited, an…
Differentiable Proximal Graph Matching
Haoru Tan, Chuang Wang, Xu-Yao Zhang +1
Graph matching is a fundamental tool in computer vision and pattern recognition. In this paper, we introduce an algorithm for graph matching based on the proximal operator, referre…
OpenMix: Exploring Outlier Samples for Misclassification Detection
Fei Zhu, Zhen Cheng, Xu-Yao Zhang +1
Reliable confidence estimation for deep neural classifiers is a challenging yet fundamental requirement in high-stakes applications. Unfortunately, modern deep neural networks are…
Unified Entropy Optimization for Open-Set Test-Time Adaptation
Zhengqing Gao, Xu-Yao Zhang, Cheng-Lin Liu
Test-time adaptation (TTA) aims at adapting a model pre-trained on the labeled source domain to the unlabeled target domain. Existing methods usually focus on improving TTA perform…
Natural Scene Character Recognition Using Robust PCA and Sparse Representation
Zheng Zhang, Yong Xu, Cheng-Lin Liu
Natural scene character recognition is challenging due to the cluttered background, which is hard to separate from text. In this paper, we propose a novel method for robust scene c…
BiNSGPS: Geometry Problem Solving via Bidirectional Neuro-Symbolic Interaction
Qi Wang, Peijie Wang, Fei Yin +1
Geometry problem solving poses distinct challenges in artificial intelligence. Existing approaches typically fall into two paradigms: symbolic methods, which exhibit limited adapta…
Drawing and Recognizing Chinese Characters with Recurrent Neural Network
Xu-Yao Zhang, Fei Yin, Yan-Ming Zhang +2
Recent deep learning based approaches have achieved great success on handwriting recognition. Chinese characters are among the most widely adopted writing systems in the world. Pre…
Online and Offline Handwritten Chinese Character Recognition: A Comprehensive Study and New Benchmark
Xu-Yao Zhang, Yoshua Bengio, Cheng-Lin Liu
Recent deep learning based methods have achieved the state-of-the-art performance for handwritten Chinese character recognition (HCCR) by learning discriminative representations di…
ProtoGCD: Unified and Unbiased Prototype Learning for Generalized Category Discovery
Shijie Ma, Fei Zhu, Xu-Yao Zhang +1
Generalized category discovery (GCD) is a pragmatic but underexplored problem, which requires models to automatically cluster and discover novel categories by leveraging the labele…
Federated Continual Instruction Tuning
Haiyang Guo, Fanhu Zeng, Fei Zhu +5
A vast amount of instruction tuning data is crucial for the impressive performance of Large Multimodal Models (LMMs), but the associated computational costs and data collection dem…
Anomaly Detection via Minimum Likelihood Generative Adversarial Networks
Chu Wang, Yan-Ming Zhang, Cheng-Lin Liu
Anomaly detection aims to detect abnormal events by a model of normality. It plays an important role in many domains such as network intrusion detection, criminal activity identity…
Open-world machine learning: A review and new outlooks
Fei Zhu, Shijie Ma, Zhen Cheng +4
Machine learning has achieved remarkable success in many applications. However, existing studies are largely based on the closed-world assumption, which assumes that the environmen…
A Lightweight Context-Driven Training-Free Network for Scene Text Segmentation and Recognition
Ritabrata Chakraborty, Shivakumara Palaiahnakote, Umapada Pal +1
Modern scene text recognition systems often depend on large end-to-end architectures that require extensive training and are prohibitively expensive for real-time scenarios. In suc…
Geoparsing: Diagram Parsing for Plane and Solid Geometry with a Unified Formal Language
Peijie Wang, Ming-Liang Zhang, Jun Cao +10
Multimodal Large Language Models (MLLMs) have achieved remarkable progress but continue to struggle with geometric reasoning, primarily due to the perception bottleneck regarding f…
Towards Non-Exemplar Semi-Supervised Class-Incremental Learning
Wenzhuo Liu, Fei Zhu, Cheng-Lin Liu
Deep neural networks perform remarkably well in close-world scenarios. However, novel classes emerged continually in real applications, making it necessary to learn incrementally.…
Towards Open-Set Text Recognition via Label-to-Prototype Learning
Chang Liu, Chun Yang, Hai-Bo Qin +3
Scene text recognition is a popular topic and extensively used in the industry. Although many methods have achieved satisfactory performance for the close-set text recognition chal…
Multi-scale Unified Network for Image Classification
Wenzhuo Liu, Fei Zhu, Cheng-Lin Liu
Convolutional Neural Networks (CNNs) have advanced significantly in visual representation learning and recognition. However, they face notable challenges in performance and computa…
Robust Classification with Convolutional Prototype Learning
Hong-Ming Yang, Xu-Yao Zhang, Fei Yin +1
Convolutional neural networks (CNNs) have been widely used for image classification. Despite its high accuracies, CNN has been shown to be easily fooled by some adversarial example…
HVR-Met: A Hypothesis-Verification-Replanning Agentic System for Extreme Weather Diagnosis
Shuo Tang, Jiadong Zhang, Gengxian Zhou +11
While deep learning-based weather forecasting paradigms have made significant strides, addressing extreme weather diagnostics remains a formidable challenge. This gap exists primar…
Active Generalized Category Discovery
Shijie Ma, Fei Zhu, Zhun Zhong +2
Generalized Category Discovery (GCD) is a pragmatic and challenging open-world task, which endeavors to cluster unlabeled samples from both novel and old classes, leveraging some l…
Scene Text Recognition with Sliding Convolutional Character Models
Fei Yin, Yi-Chao Wu, Xu-Yao Zhang +1
Scene text recognition has attracted great interests from the computer vision and pattern recognition community in recent years. State-of-the-art methods use concolutional neural n…
HiDe-LLaVA: Hierarchical Decoupling for Continual Instruction Tuning of Multimodal Large Language Model
Haiyang Guo, Fanhu Zeng, Ziwei Xiang +4
Instruction tuning is widely used to improve a pre-trained Multimodal Large Language Model (MLLM) by training it on curated task-specific datasets, enabling better comprehension of…
MSPE: Multi-Scale Patch Embedding Prompts Vision Transformers to Any Resolution
Wenzhuo Liu, Fei Zhu, Shijie Ma +1
Although Vision Transformers (ViTs) have recently advanced computer vision tasks significantly, an important real-world problem was overlooked: adapting to variable input resolutio…
Document Dewarping with Control Points
Guo-Wang Xie, Fei Yin, Xu-Yao Zhang +1
Document images are now widely captured by handheld devices such as mobile phones. The OCR performance on these images are largely affected due to geometric distortion of the docum…
A Fast Projected Fixed-Point Algorithm for Large Graph Matching
Yao Lu, Kaizhu Huang, Cheng-Lin Liu
We propose a fast approximate algorithm for large graph matching. A new projected fixed-point method is defined and a new doubly stochastic projection is adopted to derive the algo…
Recoverable Compression: A Multimodal Vision Token Recovery Mechanism Guided by Text Information
Yi Chen, Jian Xu, Xu-Yao Zhang +3
With the advancement of large-scale language modeling techniques, large multimodal models combining visual encoders with large language models have demonstrated exceptional perform…
Multimodal Agricultural Agent Architecture (MA3): A New Paradigm for Intelligent Agricultural Decision-Making
Zhuoning Xu, Jian Xu, Mingqing Zhang +3
As a strategic pillar industry for human survival and development, modern agriculture faces dual challenges: optimizing production efficiency and achieving sustainable development.…
Unified Classification and Rejection: A One-versus-All Framework
Zhen Cheng, Xu-Yao Zhang, Cheng-Lin Liu
Classifying patterns of known classes and rejecting ambiguous and novel (also called as out-of-distribution (OOD)) inputs are involved in open world pattern recognition. Deep neura…
Happy: A Debiased Learning Framework for Continual Generalized Category Discovery
Shijie Ma, Fei Zhu, Zhun Zhong +3
Constantly discovering novel concepts is crucial in evolving environments. This paper explores the underexplored task of Continual Generalized Category Discovery (C-GCD), which aim…
LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and Locating
Chao Deng, Jiale Yuan, Pi Bu +8
Large vision language models (LVLMs) have improved the document understanding capabilities remarkably, enabling the handling of complex document elements, longer contexts, and a wi…
Biologically Plausible Training of Deep Neural Networks Using a Top-down Credit Assignment Network
Jian-Hui Chen, Cheng-Lin Liu, Zuoren Wang
Despite the widespread adoption of Backpropagation algorithm-based Deep Neural Networks, the biological infeasibility of the BP algorithm could potentially limit the evolution of n…
CatalyticMLLM: A Graph-Text Multimodal Large Language Model for Catalytic Materials
Yanjie Li, Jian Xu, Xu-Yao Zhang +4
Property prediction and inverse structural design of catalytic materials are typically modeled as two independent tasks: the former predicts target properties from given structures…
Deep Direct Regression for Multi-Oriented Scene Text Detection
Wenhao He, Xu-Yao Zhang, Fei Yin +1
In this paper, we first provide a new perspective to divide existing high performance object detection methods into direct and indirect regressions. Direct regression performs boun…
Dewarping Document Image By Displacement Flow Estimation with Fully Convolutional Network
Guo-Wang Xie, Fei Yin, Xu-Yao Zhang +1
As camera-based documents are increasingly used, the rectification of distorted document images becomes a need to improve the recognition performance. In this paper, we propose a n…
PILoRA: Prototype Guided Incremental LoRA for Federated Class-Incremental Learning
Haiyang Guo, Fei Zhu, Wenzhuo Liu +2
Existing federated learning methods have effectively dealt with decentralized learning in scenarios involving data privacy and non-IID data. However, in real-world situations, each…
Practical Block-wise Neural Network Architecture Generation
Zhao Zhong, Junjie Yan, Wei Wu +2
Convolutional neural networks have gained a remarkable success in computer vision. However, most usable network architectures are hand-crafted and usually require expertise and ela…
Revisiting Confidence Estimation: Towards Reliable Failure Prediction
Fei Zhu, Xu-Yao Zhang, Zhen Cheng +1
Reliable confidence estimation is a challenging yet fundamental requirement in many risk-sensitive applications. However, modern deep neural networks are often overconfident for th…
Hybrid CNN and Dictionary-Based Models for Scene Recognition and Domain Adaptation
Guo-Sen Xie, Xu-Yao Zhang, Shuicheng Yan +1
Convolutional neural network (CNN) has achieved state-of-the-art performance in many different visual tasks. Learned from a large-scale training dataset, CNN features are much more…
SOLIDGEO: Measuring Multimodal Spatial Math Reasoning in Solid Geometry
Peijie Wang, Chao Yang, Zhong-Zhi Li +6
Geometry is a fundamental branch of mathematics and plays a crucial role in evaluating the reasoning capabilities of multimodal large language models (MLLMs). However, existing mul…
Unsupervised Structure-Texture Separation Network for Oracle Character Recognition
Mei Wang, Weihong Deng, Cheng-Lin Liu
Oracle bone script is the earliest-known Chinese writing system of the Shang dynasty and is precious to archeology and philology. However, real-world scanned oracle data are rare a…
Continual Learning for Generative AI: From LLMs to MLLMs and Beyond
Haiyang Guo, Fanhu Zeng, Fei Zhu +9
The rapid advancement of generative models has empowered modern AI systems to comprehend and produce highly sophisticated content, even achieving human-level performance in specifi…
Average of Pruning: Improving Performance and Stability of Out-of-Distribution Detection
Zhen Cheng, Fei Zhu, Xu-Yao Zhang +1
Detecting Out-of-distribution (OOD) inputs have been a critical issue for neural networks in the open world. However, the unstable behavior of OOD detection along the optimization…
Fuse, Reason and Verify: Geometry Problem Solving with Parsed Clauses from Diagram
Ming-Liang Zhang, Zhong-Zhi Li, Fei Yin +2
Geometry problem solving (GPS) requires capacities of multi-modal understanding, multi-hop reasoning and theorem knowledge application. In this paper, we propose a neural-symbolic…
Plane Geometry Diagram Parsing
Ming-Liang Zhang, Fei Yin, Yi-Han Hao +1
Geometry diagram parsing plays a key role in geometry problem solving, wherein the primitive extraction and relation parsing remain challenging due to the complex layout and betwee…
LANS: A Layout-Aware Neural Solver for Plane Geometry Problem
Zhong-Zhi Li, Ming-Liang Zhang, Fei Yin +1
Geometry problem solving (GPS) is a challenging mathematical reasoning task requiring multi-modal understanding, fusion, and reasoning. Existing neural solvers take GPS as a vision…
TL;DR: Too Long, Do Re-weighting for Efficient LLM Reasoning Compression
Zhong-Zhi Li, Xiao Liang, Zihao Tang +11
Large Language Models (LLMs) have recently achieved remarkable progress by leveraging Reinforcement Learning and extended Chain-of-Thought (CoT) techniques. However, the challenge…
Emergence of Machine Language: Towards Symbolic Intelligence with Neural Networks
Yuqi Wang, Xu-Yao Zhang, Cheng-Lin Liu +1
Representation is a core issue in artificial intelligence. Humans use discrete language to communicate and learn from each other, while machines use continuous features (like vecto…
BlockQNN: Efficient Block-wise Neural Network Architecture Generation
Zhao Zhong, Zichen Yang, Boyang Deng +4
Convolutional neural networks have gained a remarkable success in computer vision. However, most usable network architectures are hand-crafted and usually require expertise and ela…
SCAN: Sliding Convolutional Attention Network for Scene Text Recognition
Yi-Chao Wu, Fei Yin, Xu-Yao Zhang +2
Scene text recognition has drawn great attentions in the community of computer vision and artificial intelligence due to its challenges and wide applications. State-of-the-art recu…
Robust Metric Learning by Smooth Optimization
Kaizhu Huang, Rong Jin, Zenglin Xu +1
Most existing distance metric learning methods assume perfect side information that is usually given in pairwise or triplet constraints. Instead, in many real-world applications, t…
PDE-Agent: A toolchain-augmented multi-agent framework for PDE solving
Jianming Liu, Ren Zhu, Jian Xu +4
Solving Partial Differential Equations (PDEs) is a cornerstone of engineering and scientific research. Traditional methods for PDE solving are cumbersome, relying on manual setup a…
Adapting Vision-Language Models to Open Classes via Test-Time Prompt Tuning
Zhengqing Gao, Xiang Ao, Xu-Yao Zhang +1
Adapting pre-trained models to open classes is a challenging problem in machine learning. Vision-language models fully explore the knowledge of text modality, demonstrating strong…
Cross-Modal Causal Intervention for Medical Report Generation
Weixing Chen, Yang Liu, Ce Wang +4
Radiology Report Generation (RRG) is essential for computer-aided diagnosis and medication guidance, which can relieve the heavy burden of radiologists by automatically generating…
Branch-Tuning: Balancing Stability and Plasticity for Continual Self-Supervised Learning
Wenzhuo Liu, Fei Zhu, Cheng-Lin Liu
Self-supervised learning (SSL) has emerged as an effective paradigm for deriving general representations from vast amounts of unlabeled data. However, as real-world applications co…
LLaVA-c: Continual Improved Visual Instruction Tuning
Wenzhuo Liu, Fei Zhu, Haiyang Guo +2
Multimodal models like LLaVA-1.5 achieve state-of-the-art visual understanding through visual instruction tuning on multitask datasets, enabling strong instruction-following and mu…
Arbitrary Shape Scene Text Detection with Adaptive Text Region Representation
Xiaobing Wang, Yingying Jiang, Zhenbo Luo +3
Scene text detection attracts much attention in computer vision, because it can be widely used in many applications such as real-time text translation, automatic information entry,…
DocSAM: Unified Document Image Segmentation via Query Decomposition and Heterogeneous Mixed Learning
Xiao-Hui Li, Fei Yin, Cheng-Lin Liu
Document image segmentation is crucial for document analysis and recognition but remains challenging due to the diversity of document formats and segmentation tasks. Existing metho…
StylePrompter: Enhancing Domain Generalization with Test-Time Style Priors
Jiao Zhang, Jian Xu, Xu-Yao Zhang +1
In real-world applications, the sample distribution at the inference stage often differs from the one at the training stage, causing performance degradation of trained deep models.…
Ensemble Quadratic Assignment Network for Graph Matching
Haoru Tan, Chuang Wang, Sitong Wu +3
Graph matching is a commonly used technique in computer vision and pattern recognition. Recent data-driven approaches have improved the graph matching accuracy remarkably, whereas…
WPS-SAM: Towards Weakly-Supervised Part Segmentation with Foundation Models
Xinjian Wu, Ruisong Zhang, Jie Qin +2
Segmenting and recognizing diverse object parts is crucial in computer vision and robotics. Despite significant progress in object segmentation, part-level segmentation remains und…
ModalPrompt: Towards Efficient Multimodal Continual Instruction Tuning with Dual-Modality Guided Prompt
Fanhu Zeng, Fei Zhu, Haiyang Guo +2
Large Multimodal Models (LMMs) exhibit remarkable multi-tasking ability by learning mixed instruction datasets. However, novel tasks would be encountered sequentially in dynamic wo…
Weakly-Supervised Arbitrary-Shaped Text Detection with Expectation-Maximization Algorithm
Mengbiao Zhao, Wei Feng, Fei Yin +2
Arbitrary-shaped text detection is an important and challenging task in computer vision. Most existing methods require heavy data labeling efforts to produce polygon-level text reg…
CMMaTH: A Chinese Multi-modal Math Skill Evaluation Benchmark for Foundation Models
Zhong-Zhi Li, Ming-Liang Zhang, Fei Yin +7
Due to the rapid advancements in multimodal large language models, evaluating their multimodal mathematical capabilities continues to receive wide attention. Despite the datasets l…
Towards Robust Pattern Recognition: A Review
Xu-Yao Zhang, Cheng-Lin Liu, Ching Y. Suen
The accuracies for many pattern recognition tasks have increased rapidly year by year, achieving or even outperforming human performance. From the perspective of accuracy, pattern…
A Multi-Modal Neural Geometric Solver with Textual Clauses Parsed from Diagram
Ming-Liang Zhang, Fei Yin, Cheng-Lin Liu
Geometry problem solving (GPS) is a high-level mathematical reasoning requiring the capacities of multi-modal fusion and geometric knowledge application. Recently, neural solvers h…
Air Quality Prediction with A Meteorology-Guided Modality-Decoupled Spatio-Temporal Network
Hang Yin, Yan-Ming Zhang, Jian Xu +3
Air quality prediction plays a crucial role in public health and environmental protection. Accurate air quality prediction is a complex multivariate spatiotemporal problem, that in…
Towards Reliable Domain Generalization: A New Dataset and Evaluations
Jiao Zhang, Xu-Yao Zhang, Cheng-Lin Liu
There are ubiquitous distribution shifts in the real world. However, deep neural networks (DNNs) are easily biased towards the training set, which causes severe performance degrada…