Publications (63)
Paired Conditional Generative Adversarial Network for Highly Accelerated Liver 4D MRI
Di Xu, Xin Miao, Hengjie Liu +8
Purpose: 4D MRI with high spatiotemporal resolution is desired for image-guided liver radiotherapy. Acquiring densely sampling k-space data is time-consuming. Accelerated acquisiti…
Semi-supervised Counting via Pixel-by-pixel Density Distribution Modelling
Hui Lin, Zhiheng Ma, Rongrong Ji +4
This paper focuses on semi-supervised crowd counting, where only a small portion of the training data are labeled. We formulate the pixel-wise density value to regress as a probabi…
On the Use of BERT for Automated Essay Scoring: Joint Learning of Multi-Scale Essay Representation
Yongjie Wang, Chuan Wang, Ruobing Li +1
In recent years, pre-trained models have become dominant in most natural language processing (NLP) tasks. However, in the area of Automated Essay Scoring (AES), pre-trained models…
NIR-II Fluorescence Project Technology for Augmented Reality Surgical Navigation
Yuhuang Zhang, Xiaolong Liu, Zihang Liu +8
NIR-II fluorescence imaging provides superior tissue penetration and clarity, yet its clinical use in surgical navigation is hindered by a critical workflow issue. Surgeons must di…
Attention-based sequence-to-sequence model for speech recognition: development of state-of-the-art system on LibriSpeech and its application to non-native English
Yan Yin, Ramon Prieto, Bin Wang +4
Recent research has shown that attention-based sequence-to-sequence models such as Listen, Attend, and Spell (LAS) yield comparable results to state-of-the-art ASR systems on vario…
Self-Evolving Multi-Agent Framework for Efficient Decision Making in Real-Time Strategy Scenarios
Li Ma, Hao Peng, Yiming Wang +6
Large language models (LLMs) have demonstrated exceptional potential in complex reasoning,pioneering a new paradigm for autonomous agent decision making in dynamic settings. Howeve…
Remote Sensing Retrieval-Augmented Generation: Bridging Remote Sensing Imagery and Comprehensive Knowledge with a Multi-Modal Dataset and Retrieval-Augmented Generation Model
Congcong Wen, Yiting Lin, Xiaokang Qu +4
Recent progress in VLMs has demonstrated impressive capabilities across a variety of tasks in the natural image domain. Motivated by these advancements, the remote sensing communit…
A Tag Identification Approach Based On Fragile Watermark
Jianbiao Lin, Ke Ji, Hui Lin +2
This paper proposes a tag identify approach based on fragile Watermark that based on Least significant bit of the replacement that we first use a special way to initialize the cove…
SwapTalk: Audio-Driven Talking Face Generation with One-Shot Customization in Latent Space
Zeren Zhang, Haibo Qin, Jiayu Huang +4
Combining face swapping with lip synchronization technology offers a cost-effective solution for customized talking face generation. However, directly cascading existing models tog…
Brighteye: Glaucoma Screening with Color Fundus Photographs based on Vision Transformer
Hui Lin, Charilaos Apostolidis, Aggelos K. Katsaggelos
Differences in image quality, lighting conditions, and patient demographics pose challenges to automated glaucoma detection from color fundus photography. Brighteye, a method based…
DRL-STNet: Unsupervised Domain Adaptation for Cross-modality Medical Image Segmentation via Disentangled Representation Learning
Hui Lin, Florian Schiffers, Santiago López-Tapia +3
Unsupervised domain adaptation (UDA) is essential for medical image segmentation, especially in cross-modality data scenarios. UDA aims to transfer knowledge from a labeled source…
Attention-based Transducer for Online Speech Recognition
Bin Wang, Yan Yin, Hui Lin
Recent studies reveal the potential of recurrent neural network transducer (RNN-T) for end-to-end (E2E) speech recognition. Among some most popular E2E systems including RNN-T, Att…
Cosmological Constraints from Hubble parameter and SN Ia observations
Hui Lin, Cheng Hao, Xiao Wang +4
In this paper, we use a set of observational data (OHD) to constrain the CDM cosmology. This data set can be derived from the differential ages of the passively evolving…
Longitudinal Wrist PPG Analysis for Reliable Hypertension Risk Screening Using Deep Learning
Hui Lin, Jiyang Li, Ramy Hussein +6
Hypertension is a leading risk factor for cardiovascular diseases. Traditional blood pressure monitoring methods are cumbersome and inadequate for continuous tracking, prompting th…
Hybrid Ant Colony Algorithm Clonal Selection in the Application of the Cloud's Resource Scheduling
Jianbiao Lin, Yukun Zhong, Xiaowei Lin +2
In this paper, thinking over characteristics of ant colony optimization Algorithm, taking into account the characteristics of cloud computing, combined with clonal selection algori…
Learning Mixtures of Submodular Shells with Application to Document Summarization
Hui Lin, Jeff A. Bilmes
We introduce a method to learn a mixture of submodular "shells" in a large-margin setting. A submodular shell is an abstract submodular function that can be instantiated with a gro…
AdaptoNet: Modular Foundation-Adaptive Neural Networks for Cyber-Physical Attack Detection in Power Grids
Anissa Elias, Jennifer Rogers, Hui Lin +2
Cyber attacks on the power grid combine physical disruptions with compromised data to destabilize cyber-physical systems. We demonstrate that data denial attacks, where adversaries…
Flight Path Optimization with Optimal Control Method
Gaofeng Su, Xi Cheng, Siyuan Feng +5
This paper is based on a crucial issue in the aviation world: how to optimize the trajectory and controls given to the aircraft in order to optimize flight time and fuel consumptio…
DiffGAD: A Diffusion-based Unsupervised Graph Anomaly Detector
Jinghan Li, Yuan Gao, Jinda Lu +4
Graph Anomaly Detection (GAD) is crucial for identifying abnormal entities within networks, garnering significant attention across various fields. Traditional unsupervised methods,…
LGViT: Dynamic Early Exiting for Accelerating Vision Transformer
Guanyu Xu, Jiawei Hao, Li Shen +4
Recently, the efficient deployment and acceleration of powerful vision transformers (ViTs) on resource-limited edge devices for providing multimedia services have become attractive…
Structural Entropy Guided Probabilistic Coding
Xiang Huang, Hao Peng, Li Sun +4
Probabilistic embeddings have several advantages over deterministic embeddings as they map each data point to a distribution, which better describes the uncertainty and complexity…
Gramformer: Learning Crowd Counting via Graph-Modulated Transformer
Hui Lin, Zhiheng Ma, Xiaopeng Hong +2
Transformer has been popular in recent crowd counting work since it breaks the limited receptive field of traditional CNNs. However, since crowd images always contain a large numbe…
Accelerated Patient-specific Non-Cartesian MRI Reconstruction using Implicit Neural Representations
Di Xu, Hengjie Liu, Xin Miao +9
The scanning time for a fully sampled MRI can be undesirably lengthy. Compressed sensing has been developed to minimize image artifacts in accelerated scans, but the required itera…
HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation
Tianwei Lin, Wenqiao Zhang, Sijing Li +12
We present HealthGPT, a powerful Medical Large Vision-Language Model (Med-LVLM) that integrates medical visual comprehension and generation capabilities within a unified autoregres…
Smart Navigation System for Parking Assignment at Large Events: Incorporating Heterogeneous Driver Characteristics
Xi Cheng, Gaofeng Su, Siyuan Feng +5
Parking challenges escalate significantly during large events such as concerts or sports games, yet few studies address dynamic parking lot assignments for such occasions. This pap…
Automated Report-Derived Oncology VQA Benchmark for Evaluating Vision-Language Models on 3D Medical Imaging
Bo Liu, Hanxue Gu, Xiangru Li +6
Evaluating vision-language models (VLMs) on medical images requires benchmarks that are clinically grounded, scalable, and controlled for evaluation confounds. Existing public benc…
Patch Progression Masked Autoencoder with Fusion CNN Network for Classifying Evolution Between Two Pairs of 2D OCT Slices
Philippe Zhang, Weili Jiang, Yihao Li +13
Age-related Macular Degeneration (AMD) is a prevalent eye condition affecting visual acuity. Anti-vascular endothelial growth factor (anti-VEGF) treatments have been effective in s…
Hybrid CTC-Attention based End-to-End Speech Recognition using Subword Units
Zhangyu Xiao, Zhijian Ou, Wei Chu +1
In this paper, we present an end-to-end automatic speech recognition system, which successfully employs subword units in a hybrid CTC-Attention based system. The subword units are…
Object Counting: You Only Need to Look at One
Hui Lin, Xiaopeng Hong, Yabin Wang
This paper aims to tackle the challenging task of one-shot object counting. Given an image containing novel, previously unseen category objects, the goal of the task is to count al…
Cross-Silo Prototypical Calibration for Federated Learning with Non-IID Data
Zhuang Qi, Lei Meng, Zitan Chen +3
Federated Learning aims to learn a global model on the server side that generalizes to all clients in a privacy-preserving manner, by leveraging the local models from different cli…
A Simple but Effective Classification Model for Grammatical Error Correction
Zhu Kaili, Chuan Wang, Ruobing Li +3
We treat grammatical error correction (GEC) as a classification problem in this study, where for different types of errors, a target word is identified, and the classifier predicts…
MLLMRec-R1: Incentivizing Reasoning Capability in Large Language Models for Multimodal Sequential Recommendation
Yu Wang, Yonghui Yang, Le Wu +3
Group relative policy optimization (GRPO) has become a standard post-training paradigm for improving reasoning and preference alignment in large language models (LLMs), and has rec…
Airport Delay Prediction with Temporal Fusion Transformers
Ke Liu, Kaijing Ding, Xi Cheng +9
Since flight delay hurts passengers, airlines, and airports, its prediction becomes crucial for the decision-making of all stakeholders in the aviation industry and thus has been a…
Continuous Filtered Backprojection by Learnable Interpolation Network
Hui Lin, Dong Zeng, Qi Xie +3
Accurate reconstruction of computed tomography (CT) images is crucial in medical imaging field. However, there are unavoidable interpolation errors in the backprojection step of th…
Indoor Space Recognition using Deep Convolutional Neural Network: A Case Study at MIT Campus
Fan Zhang, Fabio Duarte, Ruixian Ma +3
In this paper, we propose a robust and parsimonious approach using Deep Convolutional Neural Network (DCNN) to recognize and interpret interior space. DCNN has achieved incredible…
Promote the Industry Standard of Smart Home in China by Intelligent Router Technology
Hui Lin, Jianbiao Lin, Ke Ji +2
The reason why smart home remains not popularized lies in bad product user experience, purchasing cost, and compatibility, and a lack of industry standard[1]. Echoing problems abov…
FedRSClip: Federated Learning for Remote Sensing Scene Classification Using Vision-Language Models
Hui Lin, Chao Zhang, Danfeng Hong +2
Remote sensing data is often distributed across multiple institutions, and due to privacy concerns and data-sharing restrictions, leveraging large-scale datasets in a centralized t…
Towards Real-Time Respiratory Motion Prediction based on Long Short-Term Memory Neural Networks
Hui Lin, Chengyu Shi, Brian Wang +3
Radiation therapy of thoracic and abdominal tumors requires incorporating the respiratory motion into treatments. To precisely account for the patient respiratory motions and predi…
Counterfactual Debating with Preset Stances for Hallucination Elimination of LLMs
Yi Fang, Moxin Li, Wenjie Wang +2
Large Language Models (LLMs) excel in various natural language processing tasks but struggle with hallucination issues. Existing solutions have considered utilizing LLMs' inherent…
Dance of SNN and ANN: Solving binding problem by combining spike timing and reconstructive attention
Hao Zheng, Hui Lin, Rong Zhao +1
The binding problem is one of the fundamental challenges that prevent the artificial neural network (ANNs) from a compositional understanding of the world like human perception, be…
Zero-shot Object Navigation with Vision-Language Models Reasoning
Congcong Wen, Yisiyuan Huang, Hao Huang +6
Object navigation is crucial for robots, but traditional methods require substantial training data and cannot be generalized to unknown environments. Zero-shot object navigation (Z…
MambaDSF: Multi-Scale SSM with Dilated Feature Fusion for Sonar Small Target Detection
Hui Lin, Jiayi Li, Jing Wang +1
Sonar imaging is the primary modality for underwater target detection, yet small targets remain difficult to detect due to insufficient pixel coverage, low acoustic contrast, and s…
Deep Reprogramming Distillation for Medical Foundation Models
Siyuan Du, Yuhang Zhou, Haolin Li +5
Medical foundation models pre-trained on large-scale datasets have shown powerful versatile performance. However, when adapting medical foundation models for specific medical scena…
Multi-View Attention Syntactic Enhanced Graph Convolutional Network for Aspect-based Sentiment Analysis
Xiang Huang, Hao Peng, Shuo Sun +3
Aspect-based Sentiment Analysis (ABSA) is the task aimed at predicting the sentiment polarity of aspect words within sentences. Recently, incorporating graph neural networks (GNNs)…
Polyline Path Masked Attention for Vision Transformer
Zhongchen Zhao, Chaodong Xiao, Hui Lin +3
Global dependency modeling and spatial position modeling are two core issues of the foundational architecture design in current deep learning frameworks. Recently, Vision Transform…
Boosting Crowd Counting via Multifaceted Attention
Hui Lin, Zhiheng Ma, Rongrong Ji +2
This paper focuses on the challenging crowd counting task. As large-scale variations often exist within crowd images, neither fixed-size convolution kernel of CNN nor fixed-size at…
Generalization-Enhanced Few-Shot Object Detection in Remote Sensing
Hui Lin, Nan Li, Pengjuan Yao +5
Remote sensing object detection is particularly challenging due to the high resolution, multi-scale features, and diverse ground object characteristics inherent in satellite and UA…
Metric Learning-Based Timing Synchronization by Using Lightweight Neural Network
Chaojin Qing, Na Yang, Shuhai Tang +3
Timing synchronization (TS) is one of the key tasks in orthogonal frequency division multiplexing (OFDM) systems. However, multi-path uncertainty corrupts the TS correctness, makin…
From Specificity to Generality: Revisiting Generalizable Artifacts in Detecting Face Deepfakes
Long Ma, Zhiyuan Yan, Jin Xu +5
Detecting deepfakes has been an increasingly important topic, especially given the rapid development of AI generation techniques. In this paper, we ask: How can we build a universa…
YOLO-Angio: An Algorithm for Coronary Anatomy Segmentation
Tom Liu, Hui Lin, Aggelos K. Katsaggelos +1
Coronary angiography remains the gold standard for diagnosis of coronary artery disease, the most common cause of death worldwide. While this procedure is performed more than 2 mil…
MECD: Unlocking Multi-Event Causal Discovery in Video Reasoning
Tieyuan Chen, Huabin Liu, Tianyao He +8
Video causal reasoning aims to achieve a high-level understanding of video content from a causal perspective. However, current video reasoning tasks are limited in scope, primarily…
Direct Measure Matching for Crowd Counting
Hui Lin, Xiaopeng Hong, Zhiheng Ma +4
Traditional crowd counting approaches usually use Gaussian assumption to generate pseudo density ground truth, which suffers from problems like inaccurate estimation of the Gaussia…
Experimental verification of phase discontinuities induced scintillation enhancement under weak perturbations
Han-Tao Wang, Hua-Jun Zhang, Lu Zhang +3
We verify the existence of scintillation enhancement by measuring the scintillation index of a beam composed of two coherent Gaussian vortex beams with topological charges…
Customizing Language Models with Instance-wise LoRA for Sequential Recommendation
Xiaoyu Kong, Jiancan Wu, An Zhang +4
Sequential recommendation systems predict the next interaction item based on users' past interactions, aligning recommendations with individual preferences. Leveraging the strength…
SDN-based In-network Honeypot: Preemptively Disrupt and Mislead Attacks in IoT Networks
Hui Lin
Detecting cyber attacks in the network environments used by Internet-of-things (IoT) and preventing them from causing physical perturbations play an important role in delivering de…
Semi-supervised Crowd Counting via Density Agency
Hui Lin, Zhiheng Ma, Xiaopeng Hong +2
In this paper, we propose a new agency-guided semi-supervised counting approach. First, we build a learnable auxiliary structure, namely the density agency to bring the recognized…
Rapid Reconstruction of Extremely Accelerated Liver 4D MRI via Chained Iterative Refinement
Di Xu, Xin Miao, Hengjie Liu +8
Abstract Purpose: High-quality 4D MRI requires an impractically long scanning time for dense k-space signal acquisition covering all respiratory phases. Accelerated sparse sampling…
StenUNet: Automatic Stenosis Detection from X-ray Coronary Angiography
Hui Lin, Tom Liu, Aggelos Katsaggelos +1
Coronary angiography continues to serve as the primary method for diagnosing coronary artery disease (CAD), which is the leading global cause of mortality. The severity of CAD is q…
Gated Convolutional Bidirectional Attention-based Model for Off-topic Spoken Response Detection
Yefei Zha, Ruobing Li, Hui Lin
Off-topic spoken response detection, the task aiming at predicting whether a response is off-topic for the corresponding prompt, is important for an automated speaking assessment s…
Flash EQ-Linear: Accelerating Equivariant Linear Layers via Group-wise Discrete Fourier Transform
Zhongchen Zhao, Jixin Wang, Qi Xie +4
Equivariant networks embed geometric symmetries as structural priors through weight sharing, achieving remarkable parameter efficiency across vision tasks. However, this parameter…
A knowledge transfer model for COVID-19 predicting and non-pharmaceutical intervention simulation
Jingyuan Wang, Xin Lin, Yuxi Liu +3
Since December 2019, A novel coronavirus (2019-nCoV) has been breaking out in China, which can cause respiratory diseases and severe pneumonia. Mathematical and empirical models re…
Deep Reinforcement Learning for Real-Time Ground Delay Program Revision and Corresponding Flight Delay Assignments
Ke Liu, Fan Hu, Hui Lin +6
This paper explores the optimization of Ground Delay Programs (GDP), a prevalent Traffic Management Initiative used in Air Traffic Management (ATM) to reconcile capacity and demand…
RS-MoE: A Vision-Language Model with Mixture of Experts for Remote Sensing Image Captioning and Visual Question Answering
Hui Lin, Danfeng Hong, Shuhang Ge +4
Remote Sensing Image Captioning (RSIC) presents unique challenges and plays a critical role in applications. Traditional RSIC methods often struggle to produce rich and diverse des…