papers

Publications (42)

cs.CV2021

Transparent Object Tracking Benchmark

Heng Fan, Halady Akhilesha Miththanthaya, Harshit +5

Visual tracking has achieved considerable progress in recent years. However, current research in the field mainly focuses on tracking of opaque objects, while little attention is p…

eess.SP2024

Integrated Sensing and Communication Signal Processing Based on Compressed Sensing Over Unlicensed Spectrum Bands

Haotian Liu, Zhiqing Wei, Fengyun Li +4

As a promising key technology of 6th generation (6G) mobile communication system, integrated sensing and communication (ISAC) technology aims to make full use of spectrum resources…

cond-mat.mtrl-sci2026

Substrate-Mediated Evaporation and Stochastic Evolution of Supported Au Nanoparticles

Dmitri N. Zakharov, Xiaohui Qu, Hong Wang +6

We use in situ transmission electron microscopy with automated tracking to study supported gold nanoparticles (NPs) during high-temperature vacuum annealing. \rev{The average mass…

cs.CV2025

Knowing Your Target: Target-Aware Transformer Makes Better Spatio-Temporal Video Grounding

Xin Gu, Yaojie Shen, Chenxi Luo +5

Transformer has attracted increasing interest in STVG, owing to its end-to-end pipeline and promising result. Existing Transformer-based STVG approaches often leverage a set of obj…

cs.LG2025

FM4NPP: A Scaling Foundation Model for Nuclear and Particle Physics

David Park, Shuhang Li, Yi Huang +9

Large language models have revolutionized artificial intelligence by enabling large, generalizable models trained through self-supervision. This paradigm has inspired the developme…

cs.CV2024

Efficient Temporal Action Segmentation via Boundary-aware Query Voting

Peiyao Wang, Yuewei Lin, Erik Blasch +2

Although the performance of Temporal Action Segmentation (TAS) has improved in recent years, achieving promising results often comes with a high computational cost due to dense inp…

physics.data-an2025

TPCpp-10M: Simulated proton-proton collisions in a Time Projection Chamber for AI Foundation Models

Shuhang Li, Yi Huang, David Park +10

Scientific foundation models hold great promise for advancing nuclear and particle physics by improving analysis precision and accelerating discovery. Yet, progress in this field i…

q-bio.BM2025

XDIP: A Curated X-ray Absorption Spectrum Dataset for Iron-Containing Proteins

Yufeng Wang, Peiyao Wang, Lu Wei +6

Earth-abundant iron is an essential metal in regulating the structure and function of proteins. This study presents the development of a comprehensive X-ray Absorption Spectroscopy…

cs.CV2024

AesFA: An Aesthetic Feature-Aware Arbitrary Neural Style Transfer

Joonwoo Kwon, Sooyoung Kim, Yuewei Lin +2

Neural style transfer (NST) has evolved significantly in recent years. Yet, despite its rapid progress and advancement, existing NST methods either struggle to transfer aesthetic i…

astro-ph.GA2022

Galaxy Deblending using Residual Dense Neural networks

Hong Wang, Sreevarsha Sreejith, Anže Slosar +2

We present a new neural network approach for deblending galaxy images in astronomical data using Residual Dense Neural network (RDN) architecture. We train the network on synthetic…

astro-ph.IM2023

Neural Network Based Point Spread Function Deconvolution For Astronomical Applications

Hong Wang, Sreevarsha Sreejith, Yuewei Lin +3

Optical astronomical images are strongly affected by the point spread function (PSF) of the optical system and the atmosphere (seeing) which blurs the observed image. The amount of…

cs.CV2019

Improved Deep Hashing with Soft Pairwise Similarity for Multi-label Image Retrieval

Zheng Zhang, Qin Zou, Yuewei Lin +2

Hash coding has been widely used in the approximate nearest neighbor search for large-scale image retrieval. Recently, many deep hashing methods have been proposed and shown largel…

cs.CV2026

Test-Time Registers as Global Priors for Tokenized Image Generation

Cheng-Yao Hong, Yifan Wang, Yuewei Lin +1

Attention-based models often develop attention sinks, where a small number of tokens repeatedly attract attention and accumulate unusually large activations. In vision transformers…

cs.CV2021

AGKD-BML: Defense Against Adversarial Attack by Attention Guided Knowledge Distillation and Bi-directional Metric Learning

Hong Wang, Yuefan Deng, Shinjae Yoo +2

While deep neural networks have shown impressive performance in many tasks, they are fragile to carefully designed adversarial attacks. We propose a novel adversarial training-base…

cs.CV2024

DD-RobustBench: An Adversarial Robustness Benchmark for Dataset Distillation

Yifan Wu, Jiawei Du, Ping Liu +3

Dataset distillation is an advanced technique aimed at compressing datasets into significantly smaller counterparts, while preserving formidable training performance. Significant e…

cs.CV2021

Automated Deepfake Detection

Ping Liu, Yuewei Lin, Yang He +5

In this paper, we propose to utilize Automated Machine Learning to adaptively search a neural architecture for deepfake detection. This is the first time to employ automated machin…

cs.CV2026

Together, Then Apart: Balancing Alignment and Distinctiveness for Multimodal Survival Analysis

Wenjing Liu, Qin Ren, Wen Zhang +2

The paper introduces TTA, a framework that first aligns shared patterns across histopathology images and genomic data and then preserves modality‑specific information to improve ca…

#multimodal learning#survival analysis#cancer prognosis#optimal transport
cs.CV2015

Co-interest Person Detection from Multiple Wearable Camera Videos

Yuewei Lin, Kareem Ezzeldeen, Youjie Zhou +4

Wearable cameras, such as Google Glass and Go Pro, enable video data collection over larger areas and from different views. In this paper, we tackle a new problem of locating the c…

cs.CV2023

INSURE: An Information Theory Inspired Disentanglement and Purification Model for Domain Generalization

Xi Yu, Huan-Hsin Tseng, Shinjae Yoo +2

Domain Generalization (DG) aims to learn a generalizable model on the unseen target domain by only training on the multiple observed source domains. Although a variety of DG method…

cs.CV2026

FCC: Fully Connected Correlation for One-Shot Segmentation

Seonghyeon Moon, Haein Kong, Muhammad Haris Khan +2

Few-shot segmentation (FSS) aims to segment the target object in a query image using only a small set of support images and masks. Therefore, having strong prior information for th…

cs.CV2020

Interpreting Galaxy Deblender GAN from the Discriminator's Perspective

Heyi Li, Yuewei Lin, Klaus Mueller +1

Generative adversarial networks (GANs) are well known for their unsupervised learning capabilities. A recent success in the field of astronomy is deblending two overlapping galaxy…

cs.CV2015

Unsupervised Cross-Domain Recognition by Identifying Compact Joint Subspaces

Yuewei Lin, Jing Chen, Yu Cao +4

This paper introduces a new method to solve the cross-domain recognition problem. Different from the traditional domain adaption methods which rely on a global domain shift for all…

cs.SD2026

Repurposing Image Diffusion Models for Training-Free Music Style Transfer on Mel-spectrograms

Heehwan Wang, Joonwoo Kwon, Sooyoung Kim +4

Music style transfer blends source structure with reference style to enable personalized music creation. However, existing zero-shot methods often struggle to capture fine-grained…

cs.CV2023

Defense against Adversarial Cloud Attack on Remote Sensing Salient Object Detection

Huiming Sun, Lan Fu, Jinlong Li +5

Detecting the salient objects in a remote sensing image has wide applications for the interdisciplinary research. Many existing deep learning methods have been proposed for Salient…

cs.LG2026

A Mixture of Experts Foundation Model for Scanning Electron Microscopy Image Analysis

Sk Miraj Ahmed, Yuewei Lin, Chuntian Cao +9

Scanning Electron Microscopy (SEM) is indispensable in modern materials science, enabling high-resolution imaging across a wide range of structural, chemical, and functional invest…

eess.IV2025

Macro2Micro: A Rapid and Precise Cross-modal Magnetic Resonance Imaging Synthesis using Multi-scale Structural Brain Similarity

Sooyoung Kim, Joonwoo Kwon, Junbeom Kwon +5

The human brain is a complex system requiring both macroscopic and microscopic components for comprehensive understanding. However, mapping nonlinear relationships between these sc…

cs.CV2024

LoReTrack: Efficient and Accurate Low-Resolution Transformer Tracking

Shaohua Dong, Yunhe Feng, Qing Yang +2

High-performance Transformer trackers have shown excellent results, yet they often bear a heavy computational load. Observing that a smaller input can immediately and conveniently…

cs.CV2026

Hierarchy-Guided Multimodal Representation Learning for Taxonomic Inference

Sk Miraj Ahmed, Xi Yu, Yunqi Li +2

Accurate biodiversity identification from large-scale field data is a foundational problem with direct impact on ecology, conservation, and environmental monitoring. In practice, t…

cs.CV2024

EVD4UAV: An Altitude-Sensitive Benchmark to Evade Vehicle Detection in UAV

Huiming Sun, Jiacheng Guo, Zibo Meng +4

Vehicle detection in Unmanned Aerial Vehicle (UAV) captured images has wide applications in aerial photography and remote sensing. There are many public benchmark datasets proposed…

cs.CV2015

LooseCut: Interactive Image Segmentation with Loosely Bounded Boxes

Hongkai Yu, Youjie Zhou, Hui Qian +6

One popular approach to interactively segment the foreground object of interest from an image is to annotate a bounding box that covers the foreground object. Then, a binary labeli…

cs.CV2025

VAPO: Visibility-Aware Keypoint Localization for Efficient 6DoF Object Pose Estimation

Ruyi Lian, Yuewei Lin, Longin Jan Latecki +1

Localizing predefined 3D keypoints in a 2D image is an effective way to establish 3D-2D correspondences for instance-level 6DoF object pose estimation. However, unreliable localiza…

cs.CV2025

PRVQL: Progressive Knowledge-guided Refinement for Robust Egocentric Visual Query Localization

Bing Fan, Yunhe Feng, Yapeng Tian +4

Egocentric visual query localization (EgoVQL) focuses on localizing the target of interest in space and time from first-person videos, given a visual query. Despite recent progress…

cs.CV2022

Coarse-to-fine Task-driven Inpainting for Geoscience Images

Huiming Sun, Jin Ma, Qing Guo +4

The processing and recognition of geoscience images have wide applications. Most of existing researches focus on understanding the high-quality geoscience images by assuming that a…

cs.CV2021

Point Adversarial Self Mining: A Simple Method for Facial Expression Recognition

Ping Liu, Yuewei Lin, Zibo Meng +4

In this paper, we propose a simple yet effective approach, named Point Adversarial Self Mining (PASM), to improve the recognition accuracy in facial expression recognition. Unlike…

cs.CV2023

Exploring Robust Features for Improving Adversarial Robustness

Hong Wang, Yuefan Deng, Shinjae Yoo +1

While deep neural networks (DNNs) have revolutionized many fields, their fragility to carefully designed adversarial attacks impedes the usage of DNNs in safety-critical applicatio…

cs.CV2015

Combining Local Appearance and Holistic View: Dual-Source Deep Neural Networks for Human Pose Estimation

Xiaochuan Fan, Kang Zheng, Yuewei Lin +1

We propose a new learning-based method for estimating 2D human pose from a single image, using Dual-Source Deep Convolutional Neural Networks (DS-CNN). Recently, many methods have…

cs.CV2025

Advancing from Automated to Autonomous Beamline by Leveraging Computer Vision

Baolu Li, Hongkai Yu, Huiming Sun +4

The synchrotron light source, a cutting-edge large-scale user facility, requires autonomous synchrotron beamline operations, a crucial technique that should enable experiments to b…

cs.AR2024

Empirical Measurements of AI Training Power Demand on a GPU-Accelerated Node

Imran Latif, Alex C. Newkirk, Matthew R. Carbone +5

The expansion of artificial intelligence (AI) applications has driven substantial investment in computational infrastructure, especially by cloud computing providers. Quantifying t…

cs.AI2025

Revisiting Your Memory: Reconstruction of Affect-Contextualized Memory via EEG-guided Audiovisual Generation

Joonwoo Kwon, Heehwan Wang, Jinwoo Lee +4

In this paper, we introduce RevisitAffectiveMemory, a novel task designed to reconstruct autobiographical memories through audio-visual generation guided by affect extracted from e…

cs.CV2023

RXFOOD: Plug-in RGB-X Fusion for Object of Interest Detection

Jin Ma, Jinlong Li, Qing Guo +3

The emergence of different sensors (Near-Infrared, Depth, etc.) is a remedy for the limited application scenarios of traditional RGB camera. The RGB-X tasks, which rely on RGB inpu…

q-bio.QM2026

GenCellAgent: Generalizable, Training-Free Cellular Image Segmentation via Large Language Model Agents

Xi Yu, Yang Yang, Qun Liu +3

Cellular image segmentation is essential for quantitative biology yet remains difficult due to heterogeneous modalities, morphological variability, and limited annotations. We pres…

cs.LG2025

Spectra-to-Structure and Structure-to-Spectra Inference Across the Periodic Table

Yufeng Wang, Peiyao Wang, Lu Wei +4

X-ray Absorption Spectroscopy (XAS) is a powerful technique for probing local atomic environments, yet its interpretation remains limited by the need for expert-driven analysis, co…