papers

Publications (138)

cs.CV2013

Learning Locality-Constrained Collaborative Representation for Face Recognition

Xi Peng, Lei Zhang, Zhang Yi +1

The model of low-dimensional manifold and sparse representation are two well-known concise models that suggest each data can be described by a few characteristics. Manifold learnin…

cs.CV2017

Deep Sparse Subspace Clustering

Xi Peng, Jiashi Feng, Shijie Xiao +3

In this paper, we present a deep extension of Sparse Subspace Clustering, termed Deep Sparse Subspace Clustering (DSSC). Regularized by the unit sphere distribution assumption for…

cs.CV2019

Construct Dynamic Graphs for Hand Gesture Recognition via Spatial-Temporal Attention

Yuxiao Chen, Long Zhao, Xi Peng +2

We propose a Dynamic Graph-Based Spatial-Temporal Attention (DG-STA) method for hand gesture recognition. The key idea is to first construct a fully-connected graph from a hand ske…

cs.CV2017

Reconstruction-Based Disentanglement for Pose-invariant Face Recognition

Xi Peng, Xiang Yu, Kihyuk Sohn +2

Deep neural networks (DNNs) trained on large-scale datasets have recently achieved impressive improvements in face recognition. But a persistent challenge remains to develop method…

cs.LG2025

Cross-View Graph Consistency Learning for Invariant Graph Representations

Jie Chen, Hua Mao, Wai Lok Woo +2

Graph representation learning is fundamental for analyzing graph-structured data. Exploring invariant graph representations remains a challenge for most existing graph representati…

cs.LG2019

Rethinking Kernel Methods for Node Representation Learning on Graphs

Yu Tian, Long Zhao, Xi Peng +1

Graph kernels are kernel methods measuring graph similarity and serve as a standard tool for graph classification. However, the use of kernel methods for node classification, which…

cs.CV2025

High-Speed Dynamic 3D Imaging with Sensor Fusion Splatting

Zihao Zou, Ziyuan Qu, Xi Peng +3

Capturing and reconstructing high-speed dynamic 3D scenes has numerous applications in computer graphics, vision, and interdisciplinary fields such as robotics, aerodynamics, and e…

cs.CV2023

Semantic Invariant Multi-view Clustering with Fully Incomplete Information

Pengxin Zeng, Mouxing Yang, Yiding Lu +3

Robust multi-view learning with incomplete information has received significant attention due to issues such as incomplete correspondences and incomplete instances that commonly af…

cs.CV2026

Beyond Loss Values: Robust Dynamic Pruning via Loss Trajectory Alignment

Huaiyuan Qin, Muli Yang, Gabriel James Goenawan +5

Existing dynamic data pruning methods often fail under noisy-label settings, as they typically rely on per-sample loss as the ranking criterion. This could mistakenly lead to prese…

cs.LG2024

Test-time Adaptation for Cross-modal Retrieval with Query Shift

Haobin Li, Peng Hu, Qianjun Zhang +3

The success of most existing cross-modal retrieval methods heavily relies on the assumption that the given queries follow the same distribution of the source domain. However, such…

cs.LG2026

Inside-Out: Measuring Generalization in Vision Transformers Through Inner Workings

Yunxiang Peng, Mengmeng Ma, Ziyu Yao +1

Reliable generalization metrics are fundamental to the evaluation of machine learning models. Especially in high-stakes applications where labeled target data are scarce, evaluatio…

cs.LG2024

Adaptive Cascading Network for Continual Test-Time Adaptation

Kien X. Nguyen, Fengchun Qiao, Xi Peng

We study the problem of continual test-time adaption where the goal is to adapt a source pre-trained model to a sequence of unlabelled target domains at test time. Existing methods…

cs.CV2020

Automatic Health Problem Detection from Gait Videos Using Deep Neural Networks

Rahil Mehrizi, Xi Peng, Shaoting Zhang +2

The aim of this study is developing an automatic system for detection of gait-related health problems using Deep Neural Networks (DNNs). The proposed system takes a video of patien…

cs.CV2025

LaverNet: Lightweight All-in-one Video Restoration via Selective Propagation

Haiyu Zhao, Yiwen Shan, Yuanbiao Gou +1

Recent studies have explored all-in-one video restoration, which handles multiple degradations with a unified model. However, these approaches still face two challenges when dealin…

cs.CV2018

CU-Net: Coupled U-Nets

Zhiqiang Tang, Xi Peng, Shijie Geng +2

We design a new connectivity pattern for the U-Net architecture. Given several stacked U-Nets, we couple each U-Net pair through the connections of their semantic blocks, resulting…

cs.SE2026

RuntimeSlicer: Towards Generalizable Unified Runtime State Representation for Failure Management

Lingzhe Zhang, Tong Jia, Weijie Hong +9

Modern software systems operate at unprecedented scale and complexity, where effective failure management is critical yet increasingly challenging. Metrics, traces, and logs provid…

cs.CV2025

Robust Duality Learning for Unsupervised Visible-Infrared Person Re-Identification

Yongxiang Li, Yuan Sun, Yang Qin +3

Unsupervised visible-infrared person re-identification (UVI-ReID) aims to retrieve pedestrian images across different modalities without costly annotations, but faces challenges du…

cs.AI2024

FALCON: Scalable Reasoning over Inconsistent ALC Ontologies

Tilman Hinnerichs, Zhenwei Tang, Xi Peng +2

Ontologies are one of the richest sources of knowledge. Real-world ontologies often contain thousands of axioms and are often human-made. Hence, they may contain inconsistency and…

cs.CV2026

Robust Multi-view Clustering against Imperfect Information

Zhichao Huang, Haochen Zhou, Hao Wang +2

Real-world multi-view data always suffer from imperfect information problem, where the view-specific observations are absent (i.e., Incomplete Views, IV) and cross-view corresponde…

cs.IT2015

Backhaul-Aware Caching Placement for Wireless Networks

Xi Peng, Juei-Chin Shen, Jun Zhang +1

As the capacity demand of mobile applications keeps increasing, the backhaul network is becoming a bottleneck to support high quality of experience (QoE) in next-generation wireles…

cs.LG2025

Beyond Accuracy: On the Effects of Fine-tuning Towards Vision-Language Model's Prediction Rationality

Qitong Wang, Tang Li, Kien X. Nguyen +1

Vision-Language Models (VLMs), such as CLIP, have already seen widespread applications. Researchers actively engage in further fine-tuning VLMs in safety-critical domains. In these…

cs.LG2025

Toward Robust and Harmonious Adaptation for Cross-modal Retrieval

Haobin Li, Mouxing Yang, Xi Peng

Recently, the general-to-customized paradigm has emerged as the dominant approach for Cross-Modal Retrieval (CMR), which reconciles the distribution shift problem between the sourc…

cs.SE2026

Efficient Failure Management for Multi-Agent Systems with Reasoning Trace Representation

Lingzhe Zhang, Tong Jia, Mingyu Wang +9

Large Language Models (LLM)-based Multi-Agent Systems (MASs) have emerged as a new paradigm in software system design, increasingly demonstrating strong reasoning and collaboration…

cs.CV2020

You Only Look Yourself: Unsupervised and Untrained Single Image Dehazing Neural Network

Boyun Li, Yuanbiao Gou, Shuhang Gu +3

In this paper, we study two challenging and less-touched problems in single image dehazing, namely, how to make deep learning achieve image dehazing without training on the ground-…

cs.CV2023

Non-Hierarchical Transformers for Pedestrian Segmentation

Amani Kiruga, Xi Peng

We propose a methodology to address the challenge of instance segmentation in autonomous systems, specifically targeting accessibility and inclusivity. Our approach utilizes a non-…

cs.CV2026

Learning Subspace-Preserving Sparse Attention Graphs from Heterogeneous Multiview Data

Jie Chen, Yuanbiao Gou, Chuanbin Liu +2

The high-dimensional features extracted from large-scale unlabeled data via various pretrained models with diverse architectures are referred to as heterogeneous multiview data. Mo…

cs.CV2024

Test-Time Degradation Adaptation for Open-Set Image Restoration

Yuanbiao Gou, Haiyu Zhao, Boyun Li +2

In contrast to close-set scenarios that restore images from a predefined set of degradations, open-set image restoration aims to handle the unknown degradations that were unforesee…

eess.IV2023

High-Dimensional MR Reconstruction Integrating Subspace and Adaptive Generative Models

Ruiyang Zhao, Xi Peng, Varun A. Kelkar +2

We present a novel method that integrates subspace modeling with an adaptive generative image prior for high-dimensional MR image reconstruction. The subspace model imposes an expl…

cs.LG2023

Incomplete Multi-view Clustering via Prototype-based Imputation

Haobin Li, Yunfan Li, Mouxing Yang +3

In this paper, we study how to achieve two characteristics highly-expected by incomplete multi-view clustering (IMvC). Namely, i) instance commonality refers to that within-cluster…

cs.CV2022

Are Multimodal Transformers Robust to Missing Modality?

Mengmeng Ma, Jian Ren, Long Zhao +2

Multimodal data collected from the real world are often imperfect due to missing modalities. Therefore multimodal models that are robust against modal-incomplete data are highly pr…

cs.CV2018

RED-Net: A Recurrent Encoder-Decoder Network for Video-based Face Alignment

Xi Peng, Rogerio S. Feris, Xiaoyu Wang +1

We propose a novel method for real-time face alignment in videos based on a recurrent encoder-decoder network model. Our proposed model predicts 2D facial point heat maps regulariz…

cs.LG2022

Robust Domain Adaptation for Machine Reading Comprehension

Liang Jiang, Zhenyu Huang, Jia Liu +2

Most domain adaptation methods for machine reading comprehension (MRC) use a pre-trained question-answer (QA) construction model to generate pseudo QA pairs for MRC transfer. Such…

cs.CV2019

Semantic-Guided Multi-Attention Localization for Zero-Shot Learning

Yizhe Zhu, Jianwen Xie, Zhiqiang Tang +2

Zero-shot learning extends the conventional object classification to the unseen class recognition by introducing semantic representations of classes. Existing approaches predominan…

cs.CL2026

AtlasKV: Augmenting LLMs with Billion-Scale Knowledge Graphs in 20GB VRAM

Haoyu Huang, Hong Ting Tsang, Jiaxin Bai +3

Retrieval-augmented generation (RAG) has shown some success in augmenting large language models (LLMs) with external knowledge. However, as a non-parametric knowledge integration p…

cs.LG2024

Hierarchical Sparse Representation Clustering for High-Dimensional Data Streams

Jie Chen, Hua Mao, Yuanbiao Gou +1

Data stream clustering reveals patterns within continuously arriving, potentially unbounded data sequences. Numerous data stream algorithms have been proposed to cluster data strea…

cs.CV2024

Multi-granularity Correspondence Learning from Long-term Noisy Videos

Yijie Lin, Jie Zhang, Zhenyu Huang +3

Existing video-language studies mainly focus on learning short video clips, leaving long-term temporal dependencies rarely explored due to over-high computational cost of modeling…

cs.CV2015

Constructing the L2-Graph for Robust Subspace Learning and Subspace Clustering

Xi Peng, Zhiding Yu, Huajin Tang +1

Under the framework of graph-based learning, the key to robust subspace clustering and subspace learning is to obtain a good similarity graph that eliminates the effects of errors…

cs.IT2016

Cache Size Allocation in Backhaul Limited Wireless Networks

Xi Peng, Jun Zhang, S. H. Song +1

Caching popular content at base stations is a powerful supplement to existing limited backhaul links for accommodating the exponentially increasing mobile data traffic. Given the l…

cs.CV2023

Deep learning-based estimation of whole-body kinematics from multi-view images

Kien X. Nguyen, Liying Zheng, Ashley L. Hawke +4

It is necessary to analyze the whole-body kinematics (including joint locations and joint angles) to assess risks of fatal and musculoskeletal injuries in occupational tasks. Human…

cs.LG2025

Improving Representation Learning of Complex Critical Care Data with ICU-BERT

Ricardo Santos, André V. Carreiro, Xi Peng +2

The multivariate, asynchronous nature of real-world clinical data, such as that generated in Intensive Care Units (ICUs), challenges traditional AI-based decision-support systems.…

cs.SE2026

Towards In-Depth Root Cause Localization for Microservices with Multi-Agent Recursion-of-Thought

Lingzhe Zhang, Tong Jia, Kangjin Wang +8

As modern microservice systems grow increasingly complex due to dynamic interactions and evolving runtime environments, they experience failures with rising frequency. Ensuring sys…

cs.CL2025

AutoSchemaKG: Autonomous Knowledge Graph Construction through Dynamic Schema Induction from Web-Scale Corpora

Jiaxin Bai, Wei Fan, Qi Hu +17

We present AutoSchemaKG, a framework for fully autonomous knowledge graph construction that eliminates the need for predefined schemas. Our system leverages large language models t…

cs.CV2024

A Survey on Deep Clustering: From the Prior Perspective

Yiding Lu, Haobin Li, Yunfan Li +2

Facilitated by the powerful feature extraction ability of neural networks, deep clustering has achieved great success in analyzing high-dimensional and complex real-world data. The…

cs.CV2024

SeafloorAI: A Large-scale Vision-Language Dataset for Seafloor Geological Survey

Kien X. Nguyen, Fengchun Qiao, Arthur Trembanis +1

A major obstacle to the advancements of machine learning models in marine science, particularly in sonar imagery analysis, is the scarcity of AI-ready datasets. While there have be…

math.PR2022

Tail Quantile Estimation for Non-preemptive Priority Queues

Jin Guang, Guiyu Hong, Xinyun Chen +4

Motivated by applications in computing and telecommunication systems, we investigate the problem of estimating p-quantile of steady-state sojourn times in a single-server multi-cla…

cs.LG2023

Provable Dynamic Fusion for Low-Quality Multimodal Data

Qingyang Zhang, Haitao Wu, Changqing Zhang +4

The inherent challenge of multimodal fusion is to precisely capture the cross-modal correlation and flexibly conduct cross-modal interaction. To fully release the value of each mod…

cs.CV2023

dugMatting: Decomposed-Uncertainty-Guided Matting

Jiawei Wu, Changqing Zhang, Zuoyong Li +3

Cutting out an object and estimating its opacity mask, known as image matting, is a key task in image and video editing. Due to the highly ill-posed issue, additional inputs, typic…

cs.CV2021

Uncertainty-guided Model Generalization to Unseen Domains

Fengchun Qiao, Xi Peng

We study a worst-case scenario in generalization: Out-of-domain generalization from a single source. The goal is to learn a robust model from a single source and expect it to gener…

cs.CV2019

Cartoonish sketch-based face editing in videos using identity deformation transfer

Long Zhao, Fangda Han, Xi Peng +4

We address the problem of using hand-drawn sketches to create exaggerated deformations to faces in videos, such as enlarging the shape or modifying the position of eyes or mouth. T…

eess.SP2026

Accelerated MR Elastography Using Learned Neural Network Representation

Xi Peng

To develop a deep-learning method for achieving fast high-resolution MR elastography from highly undersampled data without the need of high-quality training dataset. We first frame…

cs.CV2026

Next-Scale Prediction: A Self-Supervised Approach for Real-World Image Denoising

Yiwen Shan, Haiyu Zhao, Peng Hu +2

Self-supervised real-world image denoising remains a fundamental challenge, arising from the antagonistic trade-off between decorrelating spatially structured noise and preserving…

cs.CV2026

Robust Fuzzy Multi-view Learning under View Conflict

Siyuan Duan, Yuan Sun, Dezhong Peng +3

Trusted multi-view classification aims to deliver reliable fusion for accurate predictions and has recently attracted substantial attention in both academia and industry. However,…

cs.LG2024

Out-Of-Distribution Detection with Diversification (Provably)

Haiyun Yao, Zongbo Han, Huazhu Fu +3

Out-of-distribution (OOD) detection is crucial for ensuring reliable deployment of machine learning models. Recent advancements focus on utilizing easily accessible auxiliary outli…

cs.CV2025

PointCloud-Text Matching: Benchmark Datasets and a Baseline

Yanglin Feng, Yang Qin, Dezhong Peng +3

In this paper, we present and study a new instance-level retrieval task: PointCloud-Text Matching (PTM), which aims to identify the exact cross-modal instance that matches a given…

cs.CV2018

Quantized Densely Connected U-Nets for Efficient Landmark Localization

Zhiqiang Tang, Xi Peng, Shijie Geng +3

In this paper, we propose quantized densely connected U-Nets for efficient visual landmark localization. The idea is that features of the same semantic meanings are globally reused…

cs.CV2016

A Recurrent Encoder-Decoder Network for Sequential Face Alignment

Xi Peng, Rogerio S. Feris, Xiaoyu Wang +1

We propose a novel recurrent encoder-decoder network model for real-time video-based face alignment. Our proposed model predicts 2D facial point maps regularized by a regression lo…

cs.CV2025

Interpretable Failure Detection with Human-Level Concepts

Kien X. Nguyen, Tang Li, Xi Peng

Reliable failure detection holds paramount importance in safety-critical applications. Yet, neural networks are known to produce overconfident predictions for misclassified samples…

cs.CV2016

Connections Between Nuclear Norm and Frobenius Norm Based Representations

Xi Peng, Canyi Lu, Zhang Yi +1

A lot of works have shown that frobenius-norm based representation (FNR) is competitive to sparse representation and nuclear-norm based representation (NNR) in numerous tasks such…

cs.CV2016

Automatic Subspace Learning via Principal Coefficients Embedding

Xi Peng, Jiwen Lu, Zhang Yi +1

In this paper, we address two challenging problems in unsupervised subspace learning: 1) how to automatically identify the feature dimension of the learned subspace (i.e., automati…

cs.LG2013

Inductive Sparse Subspace Clustering

Xi Peng, Lei Zhang, Zhang Yi

Sparse Subspace Clustering (SSC) has achieved state-of-the-art clustering quality by performing spectral clustering over a -norm based similarity graph. However, SSC is a…

cs.CV2026

Multiview Self-Representation Learning across Heterogeneous Views

Jie Chen, Zhu Wang, Chuanbin Liu +1

Features of the same sample generated by different pretrained models often exhibit inherently distinct feature distributions because of discrepancies in the model pretraining objec…

cs.RO2022

Symmetry and Uncertainty-Aware Object SLAM for 6DoF Object Pose Estimation

Nathaniel Merrill, Yuliang Guo, Xingxing Zuo +5

We propose a keypoint-based object-level SLAM framework that can provide globally consistent 6DoF pose estimates for symmetric and asymmetric objects alike. To the best of our know…

cs.CV2018

A Generative Adversarial Approach for Zero-Shot Learning from Noisy Texts

Yizhe Zhu, Mohamed Elhoseiny, Bingchen Liu +2

Most existing zero-shot learning methods consider the problem as a visual semantic embedding one. Given the demonstrated capability of Generative Adversarial Networks(GANs) to gene…

cs.IT2017

Layered Group Sparse Beamforming for Cache-Enabled Green Wireless Networks

Xi Peng, Yuanming Shi, Jun Zhang +1

The exponential growth of mobile data traffic is driving the deployment of dense wireless networks, which will not only impose heavy backhaul burdens, but also generate considerabl…

cs.CV2023

Relationship Quantification of Image Degradations

Wenxin Wang, Boyun Li, Yuanbiao Gou +3

In this paper, we study two challenging but less-touched problems in image restoration, namely, i) how to quantify the relationship between image degradations and ii) how to improv…

cs.CV2020

Semantic Graph Convolutional Networks for 3D Human Pose Regression

Long Zhao, Xi Peng, Yu Tian +2

In this paper, we study the problem of learning Graph Convolutional Networks (GCNs) for regression. Current architectures of GCNs are limited to the small receptive field of convol…

cs.CL2019

Improving Distant Supervised Relation Extraction by Dynamic Neural Network

Yanjie Gou, Yinjie Lei, Lingqiao Liu +2

Distant Supervised Relation Extraction (DSRE) is usually formulated as a problem of classifying a bag of sentences that contain two query entities, into the predefined relation cla…

eess.IV2022

Multi-Scale Adaptive Network for Single Image Denoising

Yuanbiao Gou, Peng Hu, Jiancheng Lv +2

Multi-scale architectures have shown effectiveness in a variety of tasks thanks to appealing cross-scale complementarity. However, existing architectures treat different scale feat…

cs.CV2018

Learning to Forecast and Refine Residual Motion for Image-to-Video Generation

Long Zhao, Xi Peng, Yu Tian +2

We consider the problem of image-to-video translation, where an input image is translated into an output video containing motions of a single object. Recent methods for such proble…

physics.optics2021

Airy-Gaussian vortex beams in the fractional nonlinear-Schrödinger medium

Shangling He, Kangzhu Zhou, Boris A. Malomed +9

We address the propagation of vortex beams with the circular Airy-Gaussian shape in a (2+1)-dimensional optical waveguide modeled by the fractional nonlinear Schrodinger equation.…

cs.CV2024

DEAL: Disentangle and Localize Concept-level Explanations for VLMs

Tang Li, Mengmeng Ma, Xi Peng

Large pre-trained Vision-Language Models (VLMs) have become ubiquitous foundational components of other models and downstream tasks. Although powerful, our empirical results reveal…

cs.CV2026

ARK: A Dual-Axis Multimodal Retrieval Benchmark along Reasoning and Knowledge

Yijie Lin, Guofeng Ding, Haochen Zhou +3

Existing multimodal retrieval benchmarks largely emphasize semantic matching on daily-life images and offer limited diagnostics of professional knowledge and complex reasoning. To…

cs.LG2024

Beyond the Federation: Topology-aware Federated Learning for Generalization to Unseen Clients

Mengmeng Ma, Tang Li, Xi Peng

Federated Learning is widely employed to tackle distributed sensitive data. Existing methods primarily focus on addressing in-federation data heterogeneity. However, we observed th…

cs.LG2022

Twin Contrastive Learning for Online Clustering

Yunfan Li, Mouxing Yang, Dezhong Peng +3

This paper proposes to perform online clustering by conducting twin contrastive learning (TCL) at the instance and cluster level. Specifically, we find that when the data is projec…

cs.CV2026

Inside the Visual Mind: Neuroscience-Motivated Concept Circuits for Interpreting and Steering Vision Transformers

Tang Li, Yanlin Chen, Mengmeng Ma +1

Despite high accuracy, Vision Transformer (ViT) predictions can be driven by spurious cues, raising the need to understand their inner workings before safe deployment. Sparse autoe…

cs.LG2024

Beyond Accuracy: Ensuring Correct Predictions With Correct Rationales

Tang Li, Mengmeng Ma, Xi Peng

Large pretrained foundation models demonstrate exceptional performance and, in some high-stakes applications, even surpass human experts. However, most of these models are currentl…

cs.LG2025

Conditional Distribution Learning for Graph Classification

Jie Chen, Hua Mao, Chuanbin Liu +2

Leveraging the diversity and quantity of data provided by various graph-structured data augmentations while preserving intrinsic semantic information is challenging. Additionally,…

cs.LG2024

Image Clustering with External Guidance

Yunfan Li, Peng Hu, Dezhong Peng +3

The core of clustering is incorporating prior knowledge to construct supervision signals. From classic k-means based on data compactness to recent contrastive clustering guided by…

cs.AI2026

RA-CAD: Learning Post-Execution Critique for State-Aware Text-to-CAD Generation

Shuhao Yan, Changhao He, Xi Peng +1

Text-to-CAD generation translates natural-language design intent into editable and executable parametric computer-aided design (CAD) codes, reducing the expertise and effort requir…

cs.CV2025

DUDE: Diffusion-Based Unsupervised Cross-Domain Image Retrieval

Ruohong Yang, Peng Hu, Yunfan Li +1

Unsupervised cross-domain image retrieval (UCIR) aims to retrieve images of the same category across diverse domains without relying on annotations. Existing UCIR methods, which al…

cs.CV2023

Learning from Semantic Alignment between Unpaired Multiviews for Egocentric Video Recognition

Qitong Wang, Long Zhao, Liangzhe Yuan +2

We are concerned with a challenging scenario in unpaired multiview video learning. In this case, the model aims to learn comprehensive multiview representations while the cross-vie…

cs.CV2024

DiFiC: Your Diffusion Model Holds the Secret to Fine-Grained Clustering

Ruohong Yang, Peng Hu, Xi Peng +2

Fine-grained clustering is a practical yet challenging task, whose essence lies in capturing the subtle differences between instances of different classes. Such subtle differences…

cs.CV2023

Graph Matching with Bi-level Noisy Correspondence

Yijie Lin, Mouxing Yang, Jun Yu +3

In this paper, we study a novel and widely existing problem in graph matching (GM), namely, Bi-level Noisy Correspondence (BNC), which refers to node-level noisy correspondence (NN…

cs.CV2018

Jointly Optimize Data Augmentation and Network Training: Adversarial Data Augmentation in Human Pose Estimation

Xi Peng, Zhiqiang Tang, Fei Yang +2

Random data augmentation is a critical technique to avoid overfitting in training deep neural network models. However, data augmentation and network training are usually treated as…

cs.CV2022

Out-of-Domain Generalization from a Single Source: An Uncertainty Quantification Approach

Xi Peng, Fengchun Qiao, Long Zhao

We are concerned with a worst-case scenario in model generalization, in the sense that a model aims to perform well on many unseen domains while there is only one single domain ava…

cs.LG2020

Contrastive Clustering

Yunfan Li, Peng Hu, Zitao Liu +3

In this paper, we propose a one-stage online clustering method called Contrastive Clustering (CC) which explicitly performs the instance- and cluster-level contrastive learning. To…

nlin.PS2020

Stabilization of single- and multi-peak solitons in the fractional nonlinear Schroedinger equation with a trapping potential

Yunli Qiu, Boris A. Malomed, Dumitru Mihalache +3

We address the existence and stability of localized modes in the framework of the fractional nonlinear Schroedinger equation (FNSE) with the focusing cubic or focusing-defocusing c…

cs.CV2018

CR-GAN: Learning Complete Representations for Multi-view Generation

Yu Tian, Xi Peng, Long Zhao +2

Generating multi-view images from a single-view input is an essential yet challenging problem. It has broad applications in vision, graphics, and robotics. Our study indicates that…

cs.LG2015

A Unified Framework for Representation-based Subspace Clustering of Out-of-sample and Large-scale Data

Xi Peng, Huajin Tang, Lei Zhang +2

Under the framework of spectral clustering, the key of subspace clustering is building a similarity graph which describes the neighborhood relations among data points. Some recent…

cs.NI2023

dMAPAR-HMM: Reforming Traffic Model for Improving Performance Bound with Stochastic Network Calculus

Qingqing Yang, Xi Peng, Huiwen Yang +2

A popular branch of stochastic network calculus (SNC) utilizes moment-generating functions (MGFs) to characterize arrivals and services, which enables end-to-end performance analys…

cs.CV2016

Track Facial Points in Unconstrained Videos

Xi Peng, Qiong Hu, Junzhou Huang +1

Tracking Facial Points in unconstrained videos is challenging due to the non-rigid deformation that changes over time. In this paper, we propose to exploit incremental learning for…

cs.CV2025

LLaVA-ReID: Selective Multi-image Questioner for Interactive Person Re-Identification

Yiding Lu, Mouxing Yang, Dezhong Peng +3

Traditional text-based person ReID assumes that person descriptions from witnesses are complete and provided at once. However, in real-world scenarios, such descriptions are often…

eess.IV2021

Accelerated MRI Reconstruction with Separable and Enhanced Low-Rank Hankel Regularization

Xinlin Zhang, Hengfa Lu, Di Guo +5

The combination of the sparse sampling and the low-rank structured matrix reconstruction has shown promising performance, enabling a significant reduction of the magnetic resonance…

cs.LG2024

Multimodal Fusion on Low-quality Data: A Comprehensive Survey

Qingyang Zhang, Yake Wei, Zongbo Han +8

Multimodal fusion focuses on integrating information from multiple modalities with the goal of more accurate prediction, which has achieved remarkable progress in a wide range of s…

nlin.PS2020

Propagation dynamics of the circular Airy Gaussian vortex beams in the fractional nonlinear Schrödinger equation

Shangling He, Kangzhu Zhou, Xi Peng +3

We have investigated the propagation dynamics of the circular Airy Gaussian vortex beams (CAGVBs) in a (2+1)-dimesional optical system discribed by fractional nonlinear Schrödinge…

cs.LG2017

Locally linear representation for image clustering

Liangli Zhen, Zhang Yi, Xi Peng +1

It is a key to construct a similarity graph in graph-oriented subspace learning and clustering. In a similarity graph, each vertex denotes a data point and the edge weight represen…

cs.LG2019

Scalable Global Alignment Graph Kernel Using Random Features: From Node Embedding to Graph Embedding

Lingfei Wu, Ian En-Hsu Yen, Zhen Zhang +5

Graph kernels are widely used for measuring the similarity between graphs. Many existing graph kernels, which focus on local patterns within graphs rather than their global propert…

cs.CV2026

Semantic-Consistent Bidirectional Contrastive Hashing for Noisy Multi-Label Cross-Modal Retrieval

Likang Peng, Chao Su, Wenyuan Wu +4

Cross-modal hashing (CMH) facilitates efficient retrieval across different modalities (e.g., image and text) by encoding data into compact binary representations. While recent meth…

cs.LG2021

Deep Learning for Spatiotemporal Modeling of Urbanization

Tang Li, Jing Gao, Xi Peng

Urbanization has a strong impact on the health and wellbeing of populations across the world. Predictive spatial modeling of urbanization therefore can be a useful tool for effecti…

cs.CV2021

Learning View-Disentangled Human Pose Representation by Contrastive Cross-View Mutual Information Maximization

Long Zhao, Yuxiao Wang, Jiaping Zhao +7

We introduce a novel representation learning method to disentangle pose-dependent as well as view-dependent factors from 2D human poses. The method trains a network using cross-vie…