Publications (77)
Domain Adaptation with Cauchy-Schwarz Divergence
Wenzhe Yin, Shujian Yu, Yicong Lin +3
Domain adaptation aims to use training data from one or multiple source domains to learn a hypothesis that can be generalized to a different, but related, target domain. As such, h…
MvHo-IB: Multi-View Higher-Order Information Bottleneck for Brain Disorder Diagnosis
Kunyu Zhang, Qiang Li, Shujian Yu
Recent evidence suggests that modeling higher-order interactions (HOIs) in functional magnetic resonance imaging (fMRI) data can enhance the diagnostic accuracy of machine learning…
Multimodal Functional Maximum Correlation for Emotion Recognition
Deyang Zheng, Tianyi Zhang, Wenming Zheng +1
Emotional states manifest as coordinated yet heterogeneous physiological responses across central and autonomic systems, posing a fundamental challenge for multimodal representatio…
Deep Deterministic Independent Component Analysis for Hyperspectral Unmixing
Hongming Li, Shujian Yu, Jose C. Principe
We develop a new neural network based independent component analysis (ICA) method by directly minimizing the dependence amongst all extracted components. Using the matrix-based R{Ã…
Closed-Loop Adaptation for Weakly-Supervised Semantic Segmentation
Zhengqiang Zhang, Shujian Yu, Shi Yin +2
Weakly-supervised semantic segmentation aims to assign each pixel a semantic category under weak supervisions, such as image-level tags. Most of existing weakly-supervised semantic…
Multiscale Principle of Relevant Information for Hyperspectral Image Classification
Yantao Wei, Shujian Yu, Luis Sanchez Giraldo +1
This paper proposes a novel architecture, termed multiscale principle of relevant information (MPRI), to learn discriminative spectral-spatial features for hyperspectral image (HSI…
When Brain Foundation Model Meets Cauchy-Schwarz Divergence: A New Framework for Cross-Subject Motor Imagery Decoding
Jinzhou Wu, Baoping Tang, Qikang Li +3
Decoding motor imagery (MI) electroencephalogram (EEG) signals, a key non-invasive brain-computer interface (BCI) paradigm for controlling external systems, has been significantly…
BAN: Detecting Backdoors Activated by Adversarial Neuron Noise
Xiaoyun Xu, Zhuoran Liu, Stefanos Koffas +2
Backdoor attacks on deep learning represent a recent threat that has gained significant attention in the research community. Backdoor defenses are mainly based on backdoor inversio…
Coarse-to-Fine Salient Object Detection with Low-Rank Matrix Recovery
Qi Zheng, Shujian Yu, Xinge You +1
Low-Rank Matrix Recovery (LRMR) has recently been applied to saliency detection by decomposing image features into a low-rank component associated with background and a sparse comp…
Higher-order Organization in the Human Brain from Matrix-Based Rényi's Entropy
Qiang Li, Shujian Yu, Kristoffer H Madsen +2
Pairwise metrics are often employed to estimate statistical dependencies between brain regions, however they do not capture higher-order information interactions. It is critical to…
Optimal Randomized Approximations for Matrix based Renyi's Entropy
Yuxin Dong, Tieliang Gong, Shujian Yu +1
The Matrix-based Renyi's entropy enables us to directly measure information quantities from given data without the costly probability density estimation of underlying distributions…
Multi-view Hybrid Embedding: A Divide-and-Conquer Approach
Jiamiao Xu, Shujian Yu, Xinge You +3
We present a novel cross-view classification algorithm where the gallery and probe data come from different views. A popular approach to tackle this problem is the multi-view subsp…
Functional Connectome of the Human Brain with Total Correlation
Qiang Li, Greg Ver Steeg, Shujian Yu +1
Recent studies proposed the use of Total Correlation to describe functional connectivity among brain regions as a multivariate alternative to conventional pair-wise measures such a…
Explainable Multimodal Regression via Information Decomposition
Zhaozhao Ma, Shujian Yu
Multimodal regression aims to predict a continuous target from heterogeneous input sources and typically relies on fusion strategies such as early or late fusion. However, existing…
InfoDPCCA: Information-Theoretic Dynamic Probabilistic Canonical Correlation Analysis
Shiqin Tang, Shujian Yu
Extracting meaningful latent representations from high-dimensional sequential data is a crucial challenge in machine learning, with applications spanning natural science and engine…
PIME: Prototype-based Interpretable MCTS-Enhanced Brain Network Analysis for Disorder Diagnosis
Kunyu Zhang, Yanwu Yang, Jing Zhang +2
Recent deep learning methods for fMRI-based diagnosis have achieved promising accuracy by modeling functional connectivity networks. However, standard approaches often struggle wit…
HFMCA: Orthonormal Feature Learning for EEG-based Brain Decoding
Yinghao Wang, Lintao Xu, Shujian Yu +2
Electroencephalography (EEG) analysis is critical for brain-computer interfaces and neuroscience, but the intrinsic noise and high dimensionality of EEG signals hinder effective fe…
Reinforcement Learning from Cross-domain Videos with Video Prediction Model
Zhao Yang, Xinrui Zu, Jacob E. Kooi +5
Reinforcement learning from expert videos across visually distinct domains is challenging due to the absence of reward signals and the presence of domain gaps. We introduce XIPER (…
Multi-view Information Bottleneck Without Variational Approximation
Qi Zhang, Shujian Yu, Jingmin Xin +1
By "intelligently" fusing the complementary information across different views, multi-view learning is able to improve the performance of classification tasks. In this work, we ext…
Learning to Transfer with von Neumann Conditional Divergence
Ammar Shaker, Shujian Yu, Daniel Oñoro-Rubio
The similarity of feature representations plays a pivotal role in the success of problems related to domain adaptation. Feature similarity includes both the invariance of marginal…
On Kernel Method-Based Connectionist Models and Supervised Deep Learning Without Backpropagation
Shiyu Duan, Shujian Yu, Yunmei Chen +1
We propose a novel family of connectionist models based on kernel machines and consider the problem of learning layer-by-layer a compositional hypothesis class, i.e., a feedforward…
Jacobian Regularizer-based Neural Granger Causality
Wanqi Zhou, Shuanghao Bai, Shujian Yu +2
With the advancement of neural networks, diverse methods for neural Granger causality have emerged, which demonstrate proficiency in handling complex data, and nonlinear relationsh…
Learning an Interpretable Graph Structure in Multi-Task Learning
Shujian Yu, Francesco Alesiani, Ammar Shaker +1
We present a novel methodology to jointly perform multi-task learning and infer intrinsic relationship among tasks by an interpretable and sparse graph. Unlike existing multi-task…
Towards Interpretable Multi-Task Learning Using Bilevel Programming
Francesco Alesiani, Shujian Yu, Ammar Shaker +1
Interpretable Multi-Task Learning can be expressed as learning a sparse graph of the task relationship based on the prediction performance of the learned models. Since many natural…
Marine Animal Classification with Correntropy Loss Based Multi-view Learning
Zheng Cao, Shujian Yu, Bing Ouyang +4
To analyze marine animals behavior, seasonal distribution and abundance, digital imagery can be acquired by visual or Lidar camera. Depending on the quantity and properties of acqu…
Isolating Nonlinear Independent Sources in fMRI with -TCVAE Models
Qiang Li, Shujian Yu, Jesus Malo +3
Learning meaningful latent representations from nonlinear fMRI data remains a fundamental challenge in neuroimaging analysis. Traditional independent component analysis, widely use…
Information Theoretic Structured Generative Modeling
Bo Hu, Shujian Yu, Jose C. Principe
Rényi's information provides a theoretical foundation for tractable and data-efficient non-parametric density estimation, based on pair-wise evaluations in a reproducing kernel Hi…
Deep Dynamic Probabilistic Canonical Correlation Analysis
Shiqin Tang, Shujian Yu, Yining Dong +1
This paper presents Deep Dynamic Probabilistic Canonical Correlation Analysis (D2PCCA), a model that integrates deep learning with probabilistic modeling to analyze nonlinear dynam…
Multivariate Extension of Matrix-based Renyi's α-order Entropy Functional
Shujian Yu, Luis Gonzalo Sanchez Giraldo, Robert Jenssen +1
The matrix-based Renyi's α-order entropy functional was recently introduced using the normalized eigenspectrum of a Hermitian matrix of the projected data in a reproducing kernel…
OVA-IB: One vs All Information Bottleneck for Multi-Modal Alignment
Tianchao Li, Shujian Yu, Xinrui Zu +4
Contrastive learning is effective for aligning paired views or modalities, but alignment beyond two modalities remains non-trivial and comparatively underexplored. Pairwise CLIP-st…
Deep Deterministic Nonlinear ICA via Total Correlation Minimization with Matrix-Based Entropy Functional
Qiang Li, Shujian Yu, Liang Ma +4
Blind source separation, particularly through independent component analysis (ICA), is widely utilized across various signal processing domains for disentangling underlying compone…
Dual-Alignment Knowledge Retention for Continual Medical Image Segmentation
Yuxin Ye, Yan Liu, Shujian Yu
Continual learning in medical image segmentation involves sequential data acquisition across diverse domains (e.g., clinical sites), where task interference between past and curren…
Information-Theoretic Hashing for Zero-Shot Cross-Modal Retrieval
Yufeng Shi, Shujian Yu, Duanquan Xu +1
Zero-shot cross-modal retrieval (ZS-CMR) deals with the retrieval problem among heterogenous data from unseen classes. Typically, to guarantee generalization, the pre-defined class…
Aberrant High-Order Dependencies in Schizophrenia Resting-State Functional MRI Networks
Qiang Li, Vince D. Calhoun, Adithya Ram Ballem +3
The human brain has a complex, intricate functional architecture. While many studies primarily emphasize pairwise interactions, delving into high-order associations is crucial for…
Causal Recurrent Variational Autoencoder for Medical Time Series Generation
Hongming Li, Shujian Yu, Jose Principe
We propose causal recurrent variational autoencoder (CR-VAE), a novel generative model that is able to learn a Granger causal graph from a multivariate time series x and incorporat…
Robust and Fast Measure of Information via Low-rank Representation
Yuxin Dong, Tieliang Gong, Shujian Yu +2
The matrix-based Rényi's entropy allows us to directly quantify information measures from given data, without explicit estimation of the underlying probability distribution. This…
Aggregation of Dependent Expert Distributions in Multimodal Variational Autoencoders
Rogelio A Mancisidor, Robert Jenssen, Shujian Yu +1
Multimodal learning with variational autoencoders (VAEs) requires estimating joint distributions to evaluate the evidence lower bound (ELBO). Current methods, the product and mixtu…
Measuring Dependence with Matrix-based Entropy Functional
Shujian Yu, Francesco Alesiani, Xi Yu +2
Measuring the dependence of data plays a central role in statistics and machine learning. In this work, we summarize and generalize the main idea of existing information-theoretic…
Towards the Generalization of Multi-view Learning: An Information-theoretical Analysis
Wen Wen, Tieliang Gong, Yuxin Dong +2
Multiview learning has drawn widespread attention for its efficacy in leveraging cross-view consensus and complementarity information to achieve a comprehensive representation of d…
Simple stopping criteria for information theoretic feature selection
Shujian Yu, Jose C. Principe
Feature selection aims to select the smallest feature subset that yields the minimum generalization error. In the rich literature in feature selection, information theory-based app…
BrainIB++: Leveraging Graph Neural Networks and Information Bottleneck for Functional Brain Biomarkers in Schizophrenia
Tianzheng Hu, Qiang Li, Shu Liu +3
The development of diagnostic models is gaining traction in the field of psychiatric disorders. Recently, machine learning classifiers based on resting-state functional magnetic re…
Measuring the Discrepancy between Conditional Distributions: Methods, Properties and Applications
Shujian Yu, Ammar Shaker, Francesco Alesiani +1
We propose a simple yet powerful test statistic to quantify the discrepancy between two conditional distributions. The new statistic avoids the explicit estimation of the underlyin…
Modularizing Deep Learning via Pairwise Learning With Kernels
Shiyu Duan, Shujian Yu, Jose Principe
By redefining the conventional notions of layers, we present an alternative view on finitely wide, fully trainable deep neural networks as stacked linear models in feature spaces,…
Understanding Convolutional Neural Networks with Information Theory: An Initial Exploration
Shujian Yu, Kristoffer Wickstrøm, Robert Jenssen +1
The matrix-based Renyi's α-entropy functional and its multivariate extension were recently developed in terms of the normalized eigenspectrum of a Hermitian matrix of the projecte…
Robust Visual Tracking using Multi-Frame Multi-Feature Joint Modeling
Peng Zhang, Shujian Yu, Jiamiao Xu +4
It remains a huge challenge to design effective and efficient trackers under complex scenarios, including occlusions, illumination changes and pose variations. To cope with this pr…
Cross-Modal Retrieval with Cauchy-Schwarz Divergence
Jiahao Zhang, Wenzhe Yin, Shujian Yu
Effective cross-modal retrieval requires robust alignment of heterogeneous data types. Most existing methods focus on bi-modal retrieval tasks and rely on distributional alignment…
Deep Deterministic Information Bottleneck with Matrix-based Entropy Functional
Xi Yu, Shujian Yu, Jose C. Principe
We introduce the matrix-based Renyi's -order entropy functional to parameterize Tishby et al. information bottleneck (IB) principle with a neural network. We term our methodolo…
Principle of Relevant Information for Graph Sparsification
Shujian Yu, Francesco Alesiani, Wenzhe Yin +2
Graph sparsification aims to reduce the number of edges of a graph while maintaining its structural properties. In this paper, we propose the first general and effective informatio…
BrainIB: Interpretable Brain Network-based Psychiatric Diagnosis with Graph Information Bottleneck
Kaizhong Zheng, Shujian Yu, Baojuan Li +2
Developing a new diagnostic models based on the underlying biological mechanisms rather than subjective symptoms for psychiatric disorders is an emerging consensus. Recently, machi…
Understanding Autoencoders with Information Theoretic Concepts
Shujian Yu, Jose C. Principe
Despite their great success in practical applications, there is still a lack of theoretical and systematic methods to analyze deep neural networks. In this paper, we illustrate an…
Discovering Common Information in Multi-view Data
Qi Zhang, Mingfei Lu, Shujian Yu +2
We introduce an innovative and mathematically rigorous definition for computing common information from multi-view data, drawing inspiration from Gács-Körner common information i…
Efficient Brain Network Estimation with Sparse ICA in Non-Human Primate Neuroimaging
Qiang Li, Liang Ma, Masoud Seraji +4
Independent component analysis (ICA) is widely used to separate mixed signals and recover statistically independent components. However, in non-human primate neuroimaging studies,…
Modeling Higher-Order Brain Interactions via a Multi-View Information Bottleneck Framework for fMRI-based Psychiatric Diagnosis
Kunyu Zhang, Qiang Li, Vince D. Calhoun +1
Resting-state functional magnetic resonance imaging (fMRI) has emerged as a cornerstone for psychiatric diagnosis, yet most approaches rely on pairwise brain cortical or sub-cortic…
Adapting HFMCA to Graph Data: Self-Supervised Learning for Generalizable fMRI Representations
Jakub Frac, Alexander Schmatz, Qiang Li +2
Functional magnetic resonance imaging (fMRI) analysis faces significant challenges due to limited dataset sizes and domain variability between studies. Traditional self-supervised…
Cauchy-Schwarz Divergence Information Bottleneck for Regression
Shujian Yu, Xi Yu, Sigurd Løkse +2
The information bottleneck (IB) approach is popular to improve the generalization, robustness and explainability of deep neural networks. Essentially, it aims to find a minimum suf…
Gated Information Bottleneck for Generalization in Sequential Environments
Francesco Alesiani, Shujian Yu, Xi Yu
Deep neural networks suffer from poor generalization to unseen environments when the underlying data distribution is different from that in the training set. By learning minimum su…
Information Plane Analysis of Deep Neural Networks via Matrix-Based Renyi's Entropy and Tensor Kernels
Kristoffer Wickstrøm, Sigurd Løkse, Michael Kampffmeyer +3
Analyzing deep neural networks (DNNs) via information plane (IP) theory has gained tremendous attention recently as a tool to gain insight into, among others, their generalization…
Computationally Efficient Approximations for Matrix-based Renyi's Entropy
Tieliang Gong, Yuxin Dong, Shujian Yu +1
The recently developed matrix based Renyi's entropy enables measurement of information in data simply using the eigenspectrum of symmetric positive semi definite (PSD) matrices in…
Towards Uniformity and Alignment for Multimodal Representation Learning
Wenzhe Yin, Pan Zhou, Zehao Xiao +4
Multimodal representation learning aims to construct a shared embedding space in which heterogeneous modalities are semantically aligned. Despite strong empirical results, InfoNCE-…
ELEMENT: Episodic and Lifelong Exploration via Maximum Entropy
Hongming Li, Shujian Yu, Bin Liu +1
This paper proposes \emph{Episodic and Lifelong Exploration via Maximum ENTropy} (ELEMENT), a novel, multiscale, intrinsically motivated reinforcement learning (RL) framework that…
Distributional Vision-Language Alignment by Cauchy-Schwarz Divergence
Wenzhe Yin, Zehao Xiao, Pan Zhou +4
Vision-language alignment is crucial for various downstream tasks such as cross-modal generation and retrieval. Previous multimodal approaches like CLIP utilize InfoNCE to maximize…
Revisiting the Robustness of the Minimum Error Entropy Criterion: A Transfer Learning Case Study
Luis Pedro Silvestrin, Shujian Yu, Mark Hoogendoorn
Coping with distributional shifts is an important part of transfer learning methods in order to perform well in real-life tasks. However, most of the existing approaches in this ar…
R2-Trans:Fine-Grained Visual Categorization with Redundancy Reduction
Yu Wang, Shuo Ye, Shujian Yu +1
Fine-grained visual categorization (FGVC) aims to discriminate similar subcategories, whose main challenge is the large intraclass diversities and subtle inter-class differences. E…
Modular-Relatedness for Continual Learning
Ammar Shaker, Shujian Yu, Francesco Alesiani
In this paper, we propose a continual learning (CL) technique that is beneficial to sequential task learners by improving their retained accuracy and reducing catastrophic forgetti…
Interpretable Fault Detection using Projections of Mutual Information Matrix
Feiya Lv, Shujian Yu, Chenglin Wen +1
This paper presents a novel mutual information (MI) matrix based method for fault detection. Given a -dimensional fault process, the MI matrix is a matrix in which…
Bilevel Continual Learning
Ammar Shaker, Francesco Alesiani, Shujian Yu +1
Continual learning (CL) studies the problem of learning a sequence of tasks, one at a time, such that the learning of each new task does not lead to the deterioration in performanc…
PRI-VAE: Principle-of-Relevant-Information Variational Autoencoders
Yanjun Li, Shujian Yu, Jose C. Principe +2
Although substantial efforts have been made to learn disentangled representations under the variational autoencoder (VAE) framework, the fundamental properties to the dynamics of l…
Concept Drift Detection and Adaptation with Hierarchical Hypothesis Testing
Shujian Yu, Zubin Abraham, Heng Wang +3
A fundamental issue for statistical classification models in a streaming environment is that the joint distribution between predictor and response variables changes over time (a ph…
MIMIR: Masked Image Modeling for Mutual Information-based Adversarial Robustness
Xiaoyun Xu, Shujian Yu, Zhuoran Liu +1
Vision Transformers (ViTs) have emerged as a fundamental architecture and serve as the backbone of modern vision-language models. Despite their impressive performance, ViTs exhibit…
Coping with Change: Learning Invariant and Minimum Sufficient Representations for Fine-Grained Visual Categorization
Shuo Ye, Shujian Yu, Wenjin Hou +2
Fine-grained visual categorization (FGVC) is a challenging task due to similar visual appearances between various species. Previous studies always implicitly assume that the traini…
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning
Zhao Yang, Yuxuan Jiang, Ting-Chih Chen +18
Reinforcement learning (RL) has become central to LLM post-training, yet the methods that dominate current pipelines, PPO and GRPO, represent only a narrow slice of what RL offers.…
Request-and-Reverify: Hierarchical Hypothesis Testing for Concept Drift Detection with Expensive Labels
Shujian Yu, Xiaoyang Wang, Jose C. Principe
One important assumption underlying common classification models is the stationarity of the data. However, in real-world streaming applications, the data concept indicated by the j…
Continual Invariant Risk Minimization
Francesco Alesiani, Shujian Yu, Mathias Niepert
Empirical risk minimization can lead to poor generalization behavior on unseen environments if the learned model does not capture invariant feature representations. Invariant risk…
Continual Learning for fMRI-Based Brain Disorder Diagnosis via Functional Connectivity Matrices Generative Replay
Qianyu Chen, Shujian Yu
Functional magnetic resonance imaging (fMRI) is widely used for studying and diagnosing brain disorders, with functional connectivity (FC) matrices providing powerful representatio…
CI-GNN: A Granger Causality-Inspired Graph Neural Network for Interpretable Brain Network-Based Psychiatric Diagnosis
Kaizhong Zheng, Shujian Yu, Badong Chen
There is a recent trend to leverage the power of graph neural networks (GNNs) for brain-network based psychiatric diagnosis, which,in turn, also motivates an urgent need for psychi…
Generalized Cauchy-Schwarz Divergence and Its Deep Learning Applications
Mingfei Lu, Chenxu Li, Shujian Yu +2
Divergence measures play a central role and become increasingly essential in deep learning, yet efficient measures for multiple (more than two) distributions are rarely explored. T…
The Conditional Cauchy-Schwarz Divergence with Applications to Time-Series Data and Sequential Decision Making
Shujian Yu, Hongming Li, Sigurd Løkse +2
The Cauchy-Schwarz (CS) divergence was developed by PrÃncipe et al. in 2000. In this paper, we extend the classic CS divergence to quantify the closeness between two conditional d…