papers

Publications (63)

eess.IV2024

Paired Conditional Generative Adversarial Network for Highly Accelerated Liver 4D MRI

Di Xu, Xin Miao, Hengjie Liu +8

Purpose: 4D MRI with high spatiotemporal resolution is desired for image-guided liver radiotherapy. Acquiring densely sampling k-space data is time-consuming. Accelerated acquisiti…

cs.CV2024

Semi-supervised Counting via Pixel-by-pixel Density Distribution Modelling

Hui Lin, Zhiheng Ma, Rongrong Ji +4

This paper focuses on semi-supervised crowd counting, where only a small portion of the training data are labeled. We formulate the pixel-wise density value to regress as a probabi…

cs.CL2022

On the Use of BERT for Automated Essay Scoring: Joint Learning of Multi-Scale Essay Representation

Yongjie Wang, Chuan Wang, Ruobing Li +1

In recent years, pre-trained models have become dominant in most natural language processing (NLP) tasks. However, in the area of Automated Essay Scoring (AES), pre-trained models…

physics.optics2025

NIR-II Fluorescence Project Technology for Augmented Reality Surgical Navigation

Yuhuang Zhang, Xiaolong Liu, Zihang Liu +8

NIR-II fluorescence imaging provides superior tissue penetration and clarity, yet its clinical use in surgical navigation is hindered by a critical workflow issue. Surgeons must di…

cs.CL2018

Attention-based sequence-to-sequence model for speech recognition: development of state-of-the-art system on LibriSpeech and its application to non-native English

Yan Yin, Ramon Prieto, Bin Wang +4

Recent research has shown that attention-based sequence-to-sequence models such as Listen, Attend, and Spell (LAS) yield comparable results to state-of-the-art ASR systems on vario…

cs.MA2026

Self-Evolving Multi-Agent Framework for Efficient Decision Making in Real-Time Strategy Scenarios

Li Ma, Hao Peng, Yiming Wang +6

Large language models (LLMs) have demonstrated exceptional potential in complex reasoning,pioneering a new paradigm for autonomous agent decision making in dynamic settings. Howeve…

cs.CV2026

Remote Sensing Retrieval-Augmented Generation: Bridging Remote Sensing Imagery and Comprehensive Knowledge with a Multi-Modal Dataset and Retrieval-Augmented Generation Model

Congcong Wen, Yiting Lin, Xiaokang Qu +4

Recent progress in VLMs has demonstrated impressive capabilities across a variety of tasks in the natural image domain. Motivated by these advancements, the remote sensing communit…

cs.MM2014

A Tag Identification Approach Based On Fragile Watermark

Jianbiao Lin, Ke Ji, Hui Lin +2

This paper proposes a tag identify approach based on fragile Watermark that based on Least significant bit of the replacement that we first use a special way to initialize the cove…

cs.CV2024

SwapTalk: Audio-Driven Talking Face Generation with One-Shot Customization in Latent Space

Zeren Zhang, Haibo Qin, Jiayu Huang +4

Combining face swapping with lip synchronization technology offers a cost-effective solution for customized talking face generation. However, directly cascading existing models tog…

cs.CV2024

Brighteye: Glaucoma Screening with Color Fundus Photographs based on Vision Transformer

Hui Lin, Charilaos Apostolidis, Aggelos K. Katsaggelos

Differences in image quality, lighting conditions, and patient demographics pose challenges to automated glaucoma detection from color fundus photography. Brighteye, a method based…

eess.IV2024

DRL-STNet: Unsupervised Domain Adaptation for Cross-modality Medical Image Segmentation via Disentangled Representation Learning

Hui Lin, Florian Schiffers, Santiago López-Tapia +3

Unsupervised domain adaptation (UDA) is essential for medical image segmentation, especially in cross-modality data scenarios. UDA aims to transfer knowledge from a labeled source…

eess.AS2020

Attention-based Transducer for Online Speech Recognition

Bin Wang, Yan Yin, Hui Lin

Recent studies reveal the potential of recurrent neural network transducer (RNN-T) for end-to-end (E2E) speech recognition. Among some most popular E2E systems including RNN-T, Att…

astro-ph2008

Cosmological Constraints from Hubble parameter and SN Ia observations

Hui Lin, Cheng Hao, Xiao Wang +4

In this paper, we use a set of observational data (OHD) to constrain the CDM cosmology. This data set can be derived from the differential ages of the passively evolving…

eess.SP2024

Longitudinal Wrist PPG Analysis for Reliable Hypertension Risk Screening Using Deep Learning

Hui Lin, Jiyang Li, Ramy Hussein +6

Hypertension is a leading risk factor for cardiovascular diseases. Traditional blood pressure monitoring methods are cumbersome and inadequate for continuous tracking, prompting th…

cs.DC2014

Hybrid Ant Colony Algorithm Clonal Selection in the Application of the Cloud's Resource Scheduling

Jianbiao Lin, Yukun Zhong, Xiaowei Lin +2

In this paper, thinking over characteristics of ant colony optimization Algorithm, taking into account the characteristics of cloud computing, combined with clonal selection algori…

cs.LG2012

Learning Mixtures of Submodular Shells with Application to Document Summarization

Hui Lin, Jeff A. Bilmes

We introduce a method to learn a mixture of submodular "shells" in a large-margin setting. A submodular shell is an abstract submodular function that can be instantiated with a gro…

cs.CR2026

AdaptoNet: Modular Foundation-Adaptive Neural Networks for Cyber-Physical Attack Detection in Power Grids

Anissa Elias, Jennifer Rogers, Hui Lin +2

Cyber attacks on the power grid combine physical disruptions with compromised data to destabilize cyber-physical systems. We demonstrate that data denial attacks, where adversaries…

math.OC2024

Flight Path Optimization with Optimal Control Method

Gaofeng Su, Xi Cheng, Siyuan Feng +5

This paper is based on a crucial issue in the aviation world: how to optimize the trajectory and controls given to the aircraft in order to optimize flight time and fuel consumptio…

cs.LG2025

DiffGAD: A Diffusion-based Unsupervised Graph Anomaly Detector

Jinghan Li, Yuan Gao, Jinda Lu +4

Graph Anomaly Detection (GAD) is crucial for identifying abnormal entities within networks, garnering significant attention across various fields. Traditional unsupervised methods,…

cs.CV2023

LGViT: Dynamic Early Exiting for Accelerating Vision Transformer

Guanyu Xu, Jiawei Hao, Li Shen +4

Recently, the efficient deployment and acceleration of powerful vision transformers (ViTs) on resource-limited edge devices for providing multimedia services have become attractive…

cs.AI2024

Structural Entropy Guided Probabilistic Coding

Xiang Huang, Hao Peng, Li Sun +4

Probabilistic embeddings have several advantages over deterministic embeddings as they map each data point to a distribution, which better describes the uncertainty and complexity…

cs.CV2024

Gramformer: Learning Crowd Counting via Graph-Modulated Transformer

Hui Lin, Zhiheng Ma, Xiaopeng Hong +2

Transformer has been popular in recent crowd counting work since it breaks the limited receptive field of traditional CNNs. However, since crowd images always contain a large numbe…

eess.IV2025

Accelerated Patient-specific Non-Cartesian MRI Reconstruction using Implicit Neural Representations

Di Xu, Hengjie Liu, Xin Miao +9

The scanning time for a fully sampled MRI can be undesirably lengthy. Compressed sensing has been developed to minimize image artifacts in accelerated scans, but the required itera…

cs.CV2025

HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation

Tianwei Lin, Wenqiao Zhang, Sijing Li +12

We present HealthGPT, a powerful Medical Large Vision-Language Model (Med-LVLM) that integrates medical visual comprehension and generation capabilities within a unified autoregres…

cs.RO2024

Smart Navigation System for Parking Assignment at Large Events: Incorporating Heterogeneous Driver Characteristics

Xi Cheng, Gaofeng Su, Siyuan Feng +5

Parking challenges escalate significantly during large events such as concerts or sports games, yet few studies address dynamic parking lot assignments for such occasions. This pap…

cs.CV2026

Automated Report-Derived Oncology VQA Benchmark for Evaluating Vision-Language Models on 3D Medical Imaging

Bo Liu, Hanxue Gu, Xiangru Li +6

Evaluating vision-language models (VLMs) on medical images requires benchmarks that are clinically grounded, scalable, and controlled for evaluation confounds. Existing public benc…

cs.CV2025

Patch Progression Masked Autoencoder with Fusion CNN Network for Classifying Evolution Between Two Pairs of 2D OCT Slices

Philippe Zhang, Weili Jiang, Yihao Li +13

Age-related Macular Degeneration (AMD) is a prevalent eye condition affecting visual acuity. Anti-vascular endothelial growth factor (anti-VEGF) treatments have been effective in s…

eess.AS2018

Hybrid CTC-Attention based End-to-End Speech Recognition using Subword Units

Zhangyu Xiao, Zhijian Ou, Wei Chu +1

In this paper, we present an end-to-end automatic speech recognition system, which successfully employs subword units in a hybrid CTC-Attention based system. The subword units are…

cs.CV2021

Object Counting: You Only Need to Look at One

Hui Lin, Xiaopeng Hong, Yabin Wang

This paper aims to tackle the challenging task of one-shot object counting. Given an image containing novel, previously unseen category objects, the goal of the task is to count al…

cs.LG2023

Cross-Silo Prototypical Calibration for Federated Learning with Non-IID Data

Zhuang Qi, Lei Meng, Zitan Chen +3

Federated Learning aims to learn a global model on the server side that generalizes to all clients in a privacy-preserving manner, by leveraging the local models from different cli…

cs.CL2018

A Simple but Effective Classification Model for Grammatical Error Correction

Zhu Kaili, Chuan Wang, Ruobing Li +3

We treat grammatical error correction (GEC) as a classification problem in this study, where for different types of errors, a target word is identified, and the classifier predicts…

cs.IR2026

MLLMRec-R1: Incentivizing Reasoning Capability in Large Language Models for Multimodal Sequential Recommendation

Yu Wang, Yonghui Yang, Le Wu +3

Group relative policy optimization (GRPO) has become a standard post-training paradigm for improving reasoning and preference alignment in large language models (LLMs), and has rec…

cs.LG2024

Airport Delay Prediction with Temporal Fusion Transformers

Ke Liu, Kaijing Ding, Xi Cheng +9

Since flight delay hurts passengers, airlines, and airports, its prediction becomes crucial for the decision-making of all stakeholders in the aviation industry and thus has been a…

eess.IV2025

Continuous Filtered Backprojection by Learnable Interpolation Network

Hui Lin, Dong Zeng, Qi Xie +3

Accurate reconstruction of computed tomography (CT) images is crucial in medical imaging field. However, there are unavoidable interpolation errors in the backprojection step of th…

cs.CV2016

Indoor Space Recognition using Deep Convolutional Neural Network: A Case Study at MIT Campus

Fan Zhang, Fabio Duarte, Ruixian Ma +3

In this paper, we propose a robust and parsimonious approach using Deep Convolutional Neural Network (DCNN) to recognize and interpret interior space. DCNN has achieved incredible…

cs.NI2015

Promote the Industry Standard of Smart Home in China by Intelligent Router Technology

Hui Lin, Jianbiao Lin, Ke Ji +2

The reason why smart home remains not popularized lies in bad product user experience, purchasing cost, and compatibility, and a lack of industry standard[1]. Echoing problems abov…

cs.CV2025

FedRSClip: Federated Learning for Remote Sensing Scene Classification Using Vision-Language Models

Hui Lin, Chao Zhang, Danfeng Hong +2

Remote sensing data is often distributed across multiple institutions, and due to privacy concerns and data-sharing restrictions, leveraging large-scale datasets in a centralized t…

physics.med-ph2019

Towards Real-Time Respiratory Motion Prediction based on Long Short-Term Memory Neural Networks

Hui Lin, Chengyu Shi, Brian Wang +3

Radiation therapy of thoracic and abdominal tumors requires incorporating the respiratory motion into treatments. To precisely account for the patient respiratory motions and predi…

cs.CL2025

Counterfactual Debating with Preset Stances for Hallucination Elimination of LLMs

Yi Fang, Moxin Li, Wenjie Wang +2

Large Language Models (LLMs) excel in various natural language processing tasks but struggle with hallucination issues. Existing solutions have considered utilizing LLMs' inherent…

cs.AI2022

Dance of SNN and ANN: Solving binding problem by combining spike timing and reconstructive attention

Hao Zheng, Hui Lin, Rong Zhao +1

The binding problem is one of the fundamental challenges that prevent the artificial neural network (ANNs) from a compositional understanding of the world like human perception, be…

cs.RO2024

Zero-shot Object Navigation with Vision-Language Models Reasoning

Congcong Wen, Yisiyuan Huang, Hao Huang +6

Object navigation is crucial for robots, but traditional methods require substantial training data and cannot be generalized to unknown environments. Zero-shot object navigation (Z…

cs.CV2026

MambaDSF: Multi-Scale SSM with Dilated Feature Fusion for Sonar Small Target Detection

Hui Lin, Jiayi Li, Jing Wang +1

Sonar imaging is the primary modality for underwater target detection, yet small targets remain difficult to detect due to insufficient pixel coverage, low acoustic contrast, and s…

cs.CV2026

Deep Reprogramming Distillation for Medical Foundation Models

Siyuan Du, Yuhang Zhou, Haolin Li +5

Medical foundation models pre-trained on large-scale datasets have shown powerful versatile performance. However, when adapting medical foundation models for specific medical scena…

cs.CL2025

Multi-View Attention Syntactic Enhanced Graph Convolutional Network for Aspect-based Sentiment Analysis

Xiang Huang, Hao Peng, Shuo Sun +3

Aspect-based Sentiment Analysis (ABSA) is the task aimed at predicting the sentiment polarity of aspect words within sentences. Recently, incorporating graph neural networks (GNNs)…

cs.CV2025

Polyline Path Masked Attention for Vision Transformer

Zhongchen Zhao, Chaodong Xiao, Hui Lin +3

Global dependency modeling and spatial position modeling are two core issues of the foundational architecture design in current deep learning frameworks. Recently, Vision Transform…

cs.CV2022

Boosting Crowd Counting via Multifaceted Attention

Hui Lin, Zhiheng Ma, Rongrong Ji +2

This paper focuses on the challenging crowd counting task. As large-scale variations often exist within crowd images, neither fixed-size convolution kernel of CNN nor fixed-size at…

cs.CV2025

Generalization-Enhanced Few-Shot Object Detection in Remote Sensing

Hui Lin, Nan Li, Pengjuan Yao +5

Remote sensing object detection is particularly challenging due to the high resolution, multi-scale features, and diverse ground object characteristics inherent in satellite and UA…

eess.SP2023

Metric Learning-Based Timing Synchronization by Using Lightweight Neural Network

Chaojin Qing, Na Yang, Shuhai Tang +3

Timing synchronization (TS) is one of the key tasks in orthogonal frequency division multiplexing (OFDM) systems. However, multi-path uncertainty corrupts the TS correctness, makin…

cs.CV2025

From Specificity to Generality: Revisiting Generalizable Artifacts in Detecting Face Deepfakes

Long Ma, Zhiyuan Yan, Jin Xu +5

Detecting deepfakes has been an increasingly important topic, especially given the rapid development of AI generation techniques. In this paper, we ask: How can we build a universa…

eess.IV2023

YOLO-Angio: An Algorithm for Coronary Anatomy Segmentation

Tom Liu, Hui Lin, Aggelos K. Katsaggelos +1

Coronary angiography remains the gold standard for diagnosis of coronary artery disease, the most common cause of death worldwide. While this procedure is performed more than 2 mil…

cs.CV2024

MECD: Unlocking Multi-Event Causal Discovery in Video Reasoning

Tieyuan Chen, Huabin Liu, Tianyao He +8

Video causal reasoning aims to achieve a high-level understanding of video content from a causal perspective. However, current video reasoning tasks are limited in scope, primarily…

cs.CV2021

Direct Measure Matching for Crowd Counting

Hui Lin, Xiaopeng Hong, Zhiheng Ma +4

Traditional crowd counting approaches usually use Gaussian assumption to generate pseudo density ground truth, which suffers from problems like inaccurate estimation of the Gaussia…

physics.optics2021

Experimental verification of phase discontinuities induced scintillation enhancement under weak perturbations

Han-Tao Wang, Hua-Jun Zhang, Lu Zhang +3

We verify the existence of scintillation enhancement by measuring the scintillation index of a beam composed of two coherent Gaussian vortex beams with topological charges…

cs.IR2025

Customizing Language Models with Instance-wise LoRA for Sequential Recommendation

Xiaoyu Kong, Jiancan Wu, An Zhang +4

Sequential recommendation systems predict the next interaction item based on users' past interactions, aligning recommendations with individual preferences. Leveraging the strength…

cs.CR2019

SDN-based In-network Honeypot: Preemptively Disrupt and Mislead Attacks in IoT Networks

Hui Lin

Detecting cyber attacks in the network environments used by Internet-of-things (IoT) and preventing them from causing physical perturbations play an important role in delivering de…

cs.CV2022

Semi-supervised Crowd Counting via Density Agency

Hui Lin, Zhiheng Ma, Xiaopeng Hong +2

In this paper, we propose a new agency-guided semi-supervised counting approach. First, we build a learnable auxiliary structure, namely the density agency to bring the recognized…

eess.IV2024

Rapid Reconstruction of Extremely Accelerated Liver 4D MRI via Chained Iterative Refinement

Di Xu, Xin Miao, Hengjie Liu +8

Abstract Purpose: High-quality 4D MRI requires an impractically long scanning time for dense k-space signal acquisition covering all respiratory phases. Accelerated sparse sampling…

eess.IV2023

StenUNet: Automatic Stenosis Detection from X-ray Coronary Angiography

Hui Lin, Tom Liu, Aggelos Katsaggelos +1

Coronary angiography continues to serve as the primary method for diagnosing coronary artery disease (CAD), which is the leading global cause of mortality. The severity of CAD is q…

cs.CL2020

Gated Convolutional Bidirectional Attention-based Model for Off-topic Spoken Response Detection

Yefei Zha, Ruobing Li, Hui Lin

Off-topic spoken response detection, the task aiming at predicting whether a response is off-topic for the corresponding prompt, is important for an automated speaking assessment s…

cs.CV2026

Flash EQ-Linear: Accelerating Equivariant Linear Layers via Group-wise Discrete Fourier Transform

Zhongchen Zhao, Jixin Wang, Qi Xie +4

Equivariant networks embed geometric symmetries as structural priors through weight sharing, achieving remarkable parameter efficiency across vision tasks. However, this parameter…

q-bio.PE2021

A knowledge transfer model for COVID-19 predicting and non-pharmaceutical intervention simulation

Jingyuan Wang, Xin Lin, Yuxi Liu +3

Since December 2019, A novel coronavirus (2019-nCoV) has been breaking out in China, which can cause respiratory diseases and severe pneumonia. Mathematical and empirical models re…

cs.LG2024

Deep Reinforcement Learning for Real-Time Ground Delay Program Revision and Corresponding Flight Delay Assignments

Ke Liu, Fan Hu, Hui Lin +6

This paper explores the optimization of Ground Delay Programs (GDP), a prevalent Traffic Management Initiative used in Air Traffic Management (ATM) to reconcile capacity and demand…

cs.CV2025

RS-MoE: A Vision-Language Model with Mixture of Experts for Remote Sensing Image Captioning and Visual Question Answering

Hui Lin, Danfeng Hong, Shuhang Ge +4

Remote Sensing Image Captioning (RSIC) presents unique challenges and plays a critical role in applications. Traditional RSIC methods often struggle to produce rich and diverse des…