Publications (78)
Exploring Automatic Diagnosis of COVID-19 from Crowdsourced Respiratory Sound Data
Chloë Brown, Jagmohan Chauhan, Andreas Grammenos +6
Audio signals generated by the human body (e.g., sighs, breathing, heart, digestion, vibration sounds) have routinely been used by clinicians as indicators to diagnose disease or a…
Uncertainty-Aware COVID-19 Detection from Imbalanced Sound Data
Tong Xia, Jing Han, Lorena Qendro +2
Recently, sound-based COVID-19 detection studies have shown great promise to achieve scalable and prompt digital pre-screening. However, there are still two unsolved issues hinderi…
Dynamic Difficulty Awareness Training for Continuous Emotion Prediction
Zixing Zhang, Jing Han, Eduardo Coutinho +1
Time-continuous emotion prediction has become an increasingly compelling task in machine learning. Considerable efforts have been made to advance the performance of these systems.…
Learning to Refine Object Contours with a Top-Down Fully Convolutional Encoder-Decoder Network
Yahui Liu, Jian Yao, Li Li +2
We develop a novel deep contour detection algorithm with a top-down fully convolutional encoder-decoder network. Our proposed method, named TD-CEDN, solves two important issues in…
Anapole mediated giant photothermal nonlinearity in nanostructured silicon
Tianyue Zhang, Ying Che, Kai Chen +14
Featured with a plethora of electric and magnetic Mie resonances, high index dielectric nanostructures offer a versatile platform to concentrate light-matter interactions at the na…
Eigendecomposition-Based Partial FFT Demodulation for Differential OFDM in Underwater Acoustic Communications
Jing Han, Lingling Zhang, Qunfei Zhang +1
Differential orthogonal frequency division multiplexing (OFDM) is practically attractive for underwater acoustic communications since it has the potential to obviate channel estima…
Pioneering Multimodal Emotion Recognition in the Era of Large Models: From Closed Sets to Open Vocabularies
Jing Han, Zhiqiang Gao, Shihao Gao +4
Recent advances in multimodal large language models (MLLMs) have demonstrated remarkable multi- and cross-modal integration capabilities. However, their potential for fine-grained…
CAA-Net: Conditional Atrous CNNs with Attention for Explainable Device-robust Acoustic Scene Classification
Zhao Ren, Qiuqiang Kong, Jing Han +2
Acoustic Scene Classification (ASC) aims to classify the environment in which the audio signals are recorded. Recently, Convolutional Neural Networks (CNNs) have been successfully…
Tyee: A Unified, Modular, and Fully-Integrated Configurable Toolkit for Intelligent Physiological Health Care
Tao Zhou, Lingyu Shu, Zixing Zhang +1
Deep learning has shown great promise in physiological signal analysis, yet its progress is hindered by heterogeneous data formats, inconsistent preprocessing strategies, fragmente…
Affective Factors in STEM Learning and Scientific Inquiry: Assessment of Cognitive Conflict and Anxiety
Lei Bao, Yeounsoo Kim, Amy Raplinger +2
Cognitive conflict is well recognized as an important factor in conceptual change and is widely used in developing inquiry-based curricula. However, cognitive conflict can also con…
DEBATE: A Dataset for Disentangling Textual Ambiguity in Mandarin Through Speech
Haotian Guo, Jing Han, Yongfeng Tu +5
Despite extensive research on textual and visual disambiguation, disambiguation through speech (DTS) remains underexplored. This is largely due to the lack of high-quality datasets…
Fitbeat: COVID-19 Estimation based on Wristband Heart Rate
Shuo Liu, Jing Han, Estela Laporta Puyal +26
This study investigates the potential of deep learning methods to identify individuals with suspected COVID-19 infection using remotely collected heart-rate data. The study utilise…
Adversarial Training in Affective Computing and Sentiment Analysis: Recent Advances and Perspectives
Jing Han, Zixing Zhang, Nicholas Cummins +1
Over the past few years, adversarial training has become an extremely active research topic and has been successfully applied to various Artificial Intelligence (AI) domains. As a…
Dynamic 3-D measurement based on fringe-to-fringe transformation using deep learning
Haotian Yu, Xiaoyu Chen, Zhao Zhang +3
Fringe projection profilometry (FPP) has become increasingly important in dynamic 3-D shape measurement. In FPP, it is necessary to retrieve the phase of the measured object before…
A Crowdsensing Intrusion Detection Dataset For Decentralized Federated Learning Models
Chao Feng, Alberto Huertas Celdran, Jing Han +6
This paper introduces a dataset and an experimental study on Decentralized Federated Learning (DFL) for Internet of Things (IoT) crowdsensing malware detection. The dataset compris…
Towards Open Respiratory Acoustic Foundation Models: Pretraining and Benchmarking
Yuwei Zhang, Tong Xia, Jing Han +6
Respiratory audio, such as coughing and breathing sounds, has predictive power for a wide range of healthcare applications, yet is currently under-explored. The main problem for th…
MSBraM: A Multi-scale Self-supervised Brain Foundation Model for Hierarchical EEG Dynamics Learning
Tao Zhou, Jing Han, Lingyu Shu +1
Self-supervised foundation models have recently shown strong potential for electroencephalogram (EEG)-based analysis. However, existing approaches struggle to capture the inherentl…
Near-infrared Image Deblurring and Event Denoising with Synergistic Neuromorphic Imaging
Chao Qu, Shuo Zhu, Yuhang Wang +4
The fields of imaging in the nighttime dynamic and other extremely dark conditions have seen impressive and transformative advancements in recent years, partly driven by the rise o…
On Nonlocal Cohesive Continuum Mechanics and Cohesive Peridynamic Modeling (CPDM) of Inelastic Fracture
Jing Han, Shaofan Li, Haicheng Yu +2
In this work, we developed a bond-based cohesive peridynamics model (CPDM) and apply it to simulate inelastic fracture by using the meso-scale Xu-Needleman cohesive potential . By…
Fast Non-Line-of-Sight Transient Data Simulation and an Open Benchmark Dataset
Yingjie Shi, Jinye Miao, Taotao Qin +10
Non-Line-of-Sight (NLOS) imaging reconstructs the shape and depth of hidden objects from picosecond-resolved transient signals, offering potential applications in autonomous drivin…
Compositional Prototype Network with Multi-view Comparision for Few-Shot Point Cloud Semantic Segmentation
Xiaoyu Chen, Chi Zhang, Guosheng Lin +1
Point cloud segmentation is a fundamental visual understanding task in 3D vision. A fully supervised point cloud segmentation network often requires a large amount of data with poi…
SEWA DB: A Rich Database for Audio-Visual Emotion and Sentiment Research in the Wild
Jean Kossaifi, Robert Walecki, Yannis Panagakis +10
Natural human-computer interaction and audio-visual human behaviour sensing systems, which would achieve robust performance in-the-wild are more needed than ever as digital devices…
GraphMAR: Geometry-Aware Graph Learning Framework for Spatially Adaptive CT Metal Artifact Reduction
Zilong Li, Chenglong Ma, Yiming Lei +7
Computed tomography (CT) metal artifact reduction (MAR) aims to reduce the severe streaking artifacts induced by metallic implants and other high-density objects. Effective MAR gen…
Base Station Network Traffic Prediction Approach Based on LMA-DeepAR
Jiachen Zhang, Xingquan zuo, Mingying Xu +2
Accurate network traffic prediction of base station cell is very vital for the expansion and reduction of wireless devices in base station cell. The burst and uncertainty of base s…
Great Chiral Fluorescence from Optical Duality Silver Nanostructures Enabled by 3D Laser Printing
Hongjing Wen, Shichao Song, Fei Xie +9
Featured by prominent flexibility and fidelity in producing sophisticated stereoscopic structures transdimensionally, three-dimensional (3D) laser printing technique has vastly ext…
Snore-GANs: Improving Automatic Snore Sound Classification with Synthesized Data
Zixing Zhang, Jing Han, Kun Qian +3
One of the frontier issues that severely hamper the development of automatic snore sound classification (ASSC) associates to the lack of sufficient supervised training data. To cop…
Non-destructive three-dimensional measurement of hand vein based on self-supervised network
Xiaoyu Chen, Qixin Wang, Jinzhou Ge +2
At present, supervised stereo methods based on deep neural network have achieved impressive results. However, in some scenarios, accurate three-dimensional labels are inaccessible…
Exploring Automatic COVID-19 Diagnosis via voice and symptoms from Crowdsourced Data
Jing Han, Chloë Brown, Jagmohan Chauhan +6
The development of fast and accurate screening tools, which could facilitate testing and prevent more costly clinical tests, is key to the current pandemic of COVID-19. In this con…
LiDAR Data Enrichment Using Deep Learning Based on High-Resolution Image: An Approach to Achieve High-Performance LiDAR SLAM Using Low-cost LiDAR
Jiang Yue, Weisong Wen, Jing Han +1
LiDAR-based SLAM algorithms are extensively studied to providing robust and accurate positioning for autonomous driving vehicles (ADV) in the past decades. Satisfactory performance…
Sounds of COVID-19: exploring realistic performance of audio-based digital testing
Jing Han, Tong Xia, Dimitris Spathis +9
Researchers have been battling with the question of how we can identify Coronavirus disease (COVID-19) cases efficiently, affordably and at scale. Recent work has shown how audio b…
Exploring Longitudinal Cough, Breath, and Voice Data for COVID-19 Progression Prediction via Sequential Deep Learning: Model Development and Validation
Ting Dang, Jing Han, Tong Xia +9
Recent work has shown the potential of using audio data (eg, cough, breathing, and voice) in the screening for COVID-19. However, these approaches only focus on one-off detection a…
Low-Complexity Equalization of MIMO-OSDM
Jing Han, Shengqian Ma, Yujie Wang +1
Orthogonal signal-division multiplexing (OSDM) is an attractive alternative to conventional orthogonal frequency-division multiplexing (OFDM) due to its enhanced ability in peak-to…
Dual camera snapshot hyperspectral imaging system via physics informed learning
Hui Xie, Zhuang Zhao, Jing Han +3
We consider using the system's optical imaging process with convolutional neural networks (CNNs) to solve the snapshot hyperspectral imaging reconstruction problem, which uses a du…
How Structure Affects Power-Law Behavior
Jing Han, Wei Li
Complex systems contain a lot of individuals and some interactions between them. The structure of interactions can be modeled to be a network: nodes represent individuals and links…
Di- Production and Generalized Distribution Amplitudes at Future Electron-Ion Colliders
Bing'ang Guo, Jing Han, Ya-Ping Xie +1
Generalized distribution amplitudes (GDAs) offer valuable insights into the three-dimensional structure of hadrons, delineating the amplitudes associated with the transition from a…
A Summary of the ComParE COVID-19 Challenges
Harry Coppock, Alican Akman, Christian Bergler +17
The COVID-19 pandemic has caused massive humanitarian and economic damage. Teams of scientists from a broad range of disciplines have searched for methods to help governments and c…
Intelligent Cardiac Auscultation for Murmur Detection via Parallel-Attentive Models with Uncertainty Estimation
Zixing Zhang, Tao Pang, Jing Han +1
Heart murmurs are a common manifestation of cardiovascular diseases and can provide crucial clues to early cardiac abnormalities. While most current research methods primarily focu…
High Sensitivity Snapshot Spectrometer Based on Deep Network Unmixing
XiaoYu Chen, Xu Wang, Lianfa Bai +2
In this paper, we present a convolution neural network based method to recover the light intensity distribution from the overlapped dispersive spectra instead of adding an extra li…
SACodec: Asymmetric Quantization with Semantic Anchoring for Low-Bitrate High-Fidelity Neural Speech Codecs
Zhongren Dong, Bin Wang, Jing Han +4
Neural Speech Codecs face a fundamental trade-off at low bitrates: preserving acoustic fidelity often compromises semantic richness. To address this, we introduce SACodec, a novel…
From Alife Agents to a Kingdom of N Queens
Jing Han, Jiming Liu, Qingsheng Cai
This paper presents a new approach to solving N-queen problems, which involves a model of distributed autonomous agents with artificial life (ALife) and a method of representing N-…
Re-Parameterization of Lightweight Transformer for On-Device Speech Emotion Recognition
Zixing Zhang, Zhongren Dong, Weixiang Xu +1
With the increasing implementation of machine learning models on edge or Internet-of-Things (IoT) devices, deploying advanced models on resource-constrained IoT devices remains cha…
Residual Pyramid Learning for Single-Shot Semantic Segmentation
Xiaoyu Chen, Xiaotian Lou, Lianfa Bai +1
Pixel-level semantic segmentation is a challenging task with a huge amount of computation, especially if the size of input is large. In the segmentation model, apart from the featu…
ASUMOT: Motion-Consistency-Based Asynchronous UAV Detection and Tracking with Event Cameras
Baofeng Jia, Xiaoyu Chen, Jingyuan Zhang +4
The paper introduces ASUMOT, an asynchronous UAV detection and tracking framework that operates directly on raw events from event cameras by grouping motion‑consistent event blobs,…
MoRAgent: Parameter Efficient Agent Tuning with Mixture-of-Roles
Jing Han, Binwei Yan, Tianyu Guo +4
Despite recent advancements of fine-tuning large language models (LLMs) to facilitate agent tasks, parameter-efficient fine-tuning (PEFT) methodologies for agent remain largely une…
Radiologist-in-the-Loop Self-Training for Generalizable CT Metal Artifact Reduction
Chenglong Ma, Zilong Li, Yuanlin Li +5
Metal artifacts in computed tomography (CT) images can significantly degrade image quality and impede accurate diagnosis. Supervised metal artifact reduction (MAR) methods, trained…
Soft Control on Collective Behavior of a Group of Autonomous Agents by a Shill Agent
Jing Han, Ming Li, Lei Guo
This paper asks a new question: how can we control the collective behavior of self-organized multi-agent systems? We try to answer the question by proposing a new notion called 'So…
PocketLLM: Ultimate Compression of Large Language Models via Meta Networks
Ye Tian, Chengcheng Wang, Jing Han +2
As Large Language Models (LLMs) continue to grow in size, storing and transmitting them on edge devices becomes increasingly challenging. Traditional methods like quantization and…
TimeSense:Making Large Language Models Proficient in Time-Series Analysis
Zhirui Zhang, Changhua Pei, Tianyi Gao +7
In the time-series domain, an increasing number of works combine text with temporal data to leverage the reasoning capabilities of large language models (LLMs) for various downstre…
Scaling Speech Enhancement in Unseen Environments with Noise Embeddings
Gil Keren, Jing Han, Björn Schuller
We address the problem of speech enhancement generalisation to unseen environments by performing two manipulations. First, we embed an additional recording from the environment alo…
Semi-Supervised Cognitive State Classification from Speech with Multi-View Pseudo-Labeling
Yuanchao Li, Zixing Zhang, Jing Han +2
The lack of labeled data is a common challenge in speech classification tasks, particularly those requiring extensive subjective assessment, such as cognitive state classification.…
SpeCache: Speculative Key-Value Caching for Efficient Generation of LLMs
Shibo Jie, Yehui Tang, Kai Han +2
Transformer-based large language models (LLMs) have already achieved remarkable results on long-text tasks, but the limited GPU memory (VRAM) resources struggle to accommodate the…
Physics-Guided Multimodal Transformers are the Necessary Foundation for the Next Generation of Meteorological Science
Jing Han, Hanting Chen, Kai Han +4
This position paper argues that the next generation of artificial intelligence in meteorological and climate sciences must transition from fragmented hybrid heuristics toward a uni…
Structured Prompting and LLM Ensembling for Multimodal Conversational Aspect-based Sentiment Analysis
Zhiqiang Gao, Shihao Gao, Zixing Zhang +3
Understanding sentiment in multimodal conversations is a complex yet crucial challenge toward building emotionally intelligent AI systems. The Multimodal Conversational Aspect-base…
KAN-AD: Time Series Anomaly Detection with Kolmogorov-Arnold Networks
Quan Zhou, Changhua Pei, Fei Sun +6
Time series anomaly detection (TSAD) underpins real-time monitoring in cloud services and web systems, allowing rapid identification of anomalies to prevent costly failures. Most T…
Wearable Foundation Models Should Go Beyond Static Encoders
Yu Yvonne Wu, Yuwei Zhang, Hyungjun Yoon +8
Wearable foundation models (WFMs), trained on large volumes of data collected by affordable, always-on devices, have demonstrated strong performance on short-term, well-defined hea…
Accessing baryon-antibaryon generalized distribution amplitudes in
Jing Han, Bernard Pire, Qin-Tao Song
$γ^* γ\to B \bar{B}$ is the golden process to access chiral-even di-baryon generalized distribution amplitudes (GDAs) as deeply virtual Compton scattering has proven to be for th…
Revealing the Temporally Stable Bimodal Energy Distribution of FRB 20121102A with a Tripled Burst Set from AI Detections
Yidan Wang, Jing Han, Pei Wang +34
Active repeating Fast Radio Bursts (FRBs), with their large number of bursts, burst energy distribution, and their potential energy evolution, offer critical insights into the FRBs…
Benchmarking Uncertainty Quantification on Biosignal Classification Tasks under Dataset Shift
Tong Xia, Jing Han, Cecilia Mascolo
A biosignal is a signal that can be continuously measured from human bodies, such as respiratory sounds, heart activity (ECG), brain waves (EEG), etc, based on which, machine learn…
EmoBed: Strengthening Monomodal Emotion Recognition via Training with Crossmodal Emotion Embeddings
Jing Han, Zixing Zhang, Zhao Ren +1
Despite remarkable advances in emotion recognition, they are severely restrained from either the essentially limited property of the employed single modality, or the synchronous pr…
Baryon-antibaryon generalized distribution amplitudes and
Jing Han, Bernard Pire, Qin-Tao Song
Baryon-antibaryon generalized distribution amplitudes (GDAs) give an access to timelike gravitational form factors (GFFs) which are complementary to the spacelike ones which can be…
Segmentation-free Heart Pathology Detection Using Deep Learning
Erika Bondareva, Jing Han, William Bradlow +1
Cardiovascular (CV) diseases are the leading cause of death in the world, and auscultation is typically an essential part of a cardiovascular examination. The ability to diagnose a…
DiC: Rethinking Conv3x3 Designs in Diffusion Models
Yuchuan Tian, Jing Han, Chengcheng Wang +3
Diffusion models have shown exceptional performance in visual generation tasks. Recently, these models have shifted from traditional U-Shaped CNN-Attention hybrid structures to ful…
A Reinforcement Learning-based Adaptive Control Model for Future Street Planning, An Algorithm and A Case Study
Qiming Ye, Yuxiang Feng, Jing Han +2
With the emerging technologies in Intelligent Transportation System (ITS), the adaptive operation of road space is likely to be realised within decades. An intelligent street can l…
High-SNR snapshot multiplex spectrometer with sub-Hadamard-S matrix coding
Zhuang Zhao, Lianfa Bai, Jing Han +1
We present a robust high signal-to-noise ratio (SNR) snapshot multiplex spectrometer with sub-Hadamard-S matrix coding. We demonstrated for the first time that the sub-Hadamard-S m…
Refashioning Emotion Recognition Modelling: The Advent of Generalised Large Models
Zixing Zhang, Liyizhe Peng, Tao Pang +3
After the inception of emotion recognition or affective computing, it has increasingly become an active research topic due to its broad applications. Over the past couple of decade…
Synchronous locating and imaging behind scattering medium in a large depth based on deep learning
Shuo Zhu, Enlai Guo, Qianying Cui +3
Scattering medium brings great difficulties to locate and image planar objects especially when the object has a large depth. In this letter, a novel learning-based method is presen…
Learning-based real-time method to looking through scattering medium beyond the memory effect
Enlai Guo, Shuo Zhu, Yan Sun +2
Strong scattering medium brings great difficulties to optical imaging, which is also a problem in medical imaging and many other fields. Optical memory effect makes it possible to…
Learning audio sequence representations for acoustic event classification
Zixing Zhang, Ding Liu, Jing Han +2
Acoustic Event Classification (AEC) has become a significant task for machines to perceive the surrounding auditory scene. However, extracting effective representations that captur…
Post-Training Quantization for Diffusion Transformer via Hierarchical Timestep Grouping
Ning Ding, Jing Han, Yuchuan Tian +3
Diffusion Transformer (DiT) has now become the preferred choice for building image generation models due to its great generation capability. Unlike previous convolution-based UNet…
Learning of Content Knowledge and Development of Scientific Reasoning Ability: A Cross Culture Comparison
Lei Bao, Kai Fang, Tianfang Cai +6
Student content knowledge and general reasoning abilities are two important areas in education practice and research. However, there hasn't been much work in physics education that…
AudioFab: Building A General and Intelligent Audio Factory through Tool Learning
Cheng Zhu, Jing Han, Qianshuai Xue +3
Currently, artificial intelligence is profoundly transforming the audio domain; however, numerous advanced algorithms and tools remain fragmented, lacking a unified and efficient f…
HAFFormer: A Hierarchical Attention-Free Framework for Alzheimer's Disease Detection From Spontaneous Speech
Zhongren Dong, Zixing Zhang, Weixiang Xu +3
Automatically detecting Alzheimer's Disease (AD) from spontaneous speech plays an important role in its early diagnosis. Recent approaches highly rely on the Transformer architectu…
Cross-device Federated Learning for Mobile Health Diagnostics: A First Study on COVID-19 Detection
Tong Xia, Jing Han, Abhirup Ghosh +1
Federated learning (FL) aided health diagnostic models can incorporate data from a large number of personal edge devices (e.g., mobile phones) while keeping the data local to the o…
Indoor simultaneous localization and mapping based on fringe projection profilometry
Yang Zhao, Kai Zhang, Haotian Yu +3
Simultaneous Localization and Mapping (SLAM) plays an important role in outdoor and indoor applications ranging from autonomous driving to indoor robotics. Outdoor SLAM has been wi…
The INTERSPEECH 2021 Computational Paralinguistics Challenge: COVID-19 Cough, COVID-19 Speech, Escalation & Primates
Björn W. Schuller, Anton Batliner, Christian Bergler +21
The INTERSPEECH 2021 Computational Paralinguistics Challenge addresses four different problems for the first time in a research competition under well-defined conditions: In the CO…
An Early Study on Intelligent Analysis of Speech under COVID-19: Severity, Sleep Quality, Fatigue, and Anxiety
Jing Han, Kun Qian, Meishu Song +11
The COVID-19 outbreak was announced as a global pandemic by the World Health Organisation in March 2020 and has affected a growing number of people in the past few weeks. In this c…
Customising General Large Language Models for Specialised Emotion Recognition Tasks
Liyizhe Peng, Zixing Zhang, Tao Pang +4
The advent of large language models (LLMs) has gained tremendous attention over the past year. Previous studies have shown the astonishing performance of LLMs not only in other tas…
TimeSeriesBench: An Industrial-Grade Benchmark for Time Series Anomaly Detection Models
Haotian Si, Jianhui Li, Changhua Pei +9
Time series anomaly detection (TSAD) has gained significant attention due to its real-world applications to improve the stability of modern software systems. However, there is no e…