papers

Publications (74)

cond-mat.mtrl-sci2024

Electrical Control Grain Dimensionality with Multilevel Magnetic Anisotropy

Shengyao Li, Sabpreet Bhatti, Siew Lang Teo +11

In alignment with the increasing demand for larger storage capacity and longer data retention, electrical control of magnetic anisotropy has been a research focus in the realm of s…

cs.CV2023

DeepMAD: Mathematical Architecture Design for Deep Convolutional Neural Network

Xuan Shen, Yaohua Wang, Ming Lin +4

The rapid advances in Vision Transformer (ViT) refresh the state-of-the-art performances in various vision tasks, overshadowing the conventional CNN-based models. This ignites a fe…

cond-mat.mtrl-sci2024

Towards edge engineering of two-dimensional layered transition-metal dichalcogenides by chemical vapor deposition

Wei Fu, Mark John, Thathsara D. Maddumapatabandi +4

The manipulation of edge configurations and structures in atomically thin transition metal dichalcogenides (TMDs) for versatile functionalization has attracted intensive interest i…

cond-mat.mtrl-sci2025

Data-Driven Design-Test-Make-Analyze Paradigm for Inorganic Crystals: Ultrafast Synthesis of Ternary Oxides

Haiwen Dai, Matthew J. McDermott, Andy Paul Chen +19

Data-driven methodologies hold the promise of revolutionizing inorganic materials discovery, but they often face challenges due to discrepancies between theoretical predictions and…

cs.CV2021

Improving Generalization of Transfer Learning Across Domains Using Spatio-Temporal Features in Autonomous Driving

Shivam Akhauri, Laura Zheng, Tom Goldstein +1

Practical learning-based autonomous driving models must be capable of generalizing learned behaviors from simulated to real domains, and from training data to unseen domains with u…

stat.ML2017

The Second Order Linear Model

Ming Lin, Shuang Qiu, Bin Hong +1

We study a fundamental class of regression models called the second order linear model (SLM). The SLM extends the linear model to high order functional space and has attracted cons…

eess.IV2022

Learning Accurate Entropy Model with Global Reference for Image Compression

Yichen Qian, Zhiyu Tan, Xiuyu Sun +5

In recent deep image compression neural networks, the entropy model plays a critical role in estimating the prior distribution of deep image encodings. Existing methods combine hyp…

physics.app-ph2022

Synaptic modulation of conductivity and magnetism in a CoPt-based electrochemical transistor

Shengyao Li, Bojun Miao, Xueyan Wang +5

Among various types of neuromorphic devices towards artificial intelligence, the electrochemical synaptic transistor emerges, in which the channel conductance is modulated by the i…

stat.ML2019

Which Factorization Machine Modeling is Better: A Theoretical Answer with Optimal Guarantee

Ming Lin, Shuang Qiu, Jieping Ye +5

Factorization machine (FM) is a popular machine learning model to capture the second order feature interactions. The optimal learning guarantee of FM and its generalized version is…

cs.LG2025

Model Steering: Learning with a Reference Model Improves Generalization Bounds and Scaling Laws

Xiyuan Wei, Ming Lin, Fanjiang Ye +4

This paper formalizes an emerging learning paradigm that uses a trained model as a reference to guide and enhance the training of a target model through strategic data selection or…

cs.RO2023

GeoLCR: Attention-based Geometric Loop Closure and Registration

Jing Liang, Sanghyun Son, Ming Lin +1

We present a novel algorithm specially designed for loop detection and registration that utilizes Lidar-based perception. Our approach to loop detection involves voxelizing point c…

stat.ML2019

Robust Gaussian Process Regression for Real-Time High Precision GPS Signal Enhancement

Ming Lin, Xiaomin Song, Qi Qian +4

Satellite-based positioning system such as GPS often suffers from large amount of noise that degrades the positioning accuracy dramatically especially in real-time applications. In…

cs.CL2026

DRTriton: Large-Scale Synthetic Data Driven Reinforcement Learning for Triton Kernel Generation

Siqi Guo, Ming Lin, Tianbao Yang

Developing efficient CUDA kernels is a fundamental yet challenging task in the generative AI industry. Recent research leverages Large Language Models (LLMs) to automatically conve…

cs.CV2023

Maximizing Spatio-Temporal Entropy of Deep 3D CNNs for Efficient Video Recognition

Junyan Wang, Zhenhong Sun, Yichen Qian +5

3D convolution neural networks (CNNs) have been the prevailing option for video recognition. To capture the temporal information, 3D convolutions are computed along the sequences,…

cs.CV2020

Neural Architecture Design for GPU-Efficient Networks

Ming Lin, Hesen Chen, Xiuyu Sun +3

Many mission-critical systems are based on GPU for inference. It requires not only high recognition accuracy but also low latency in responding time. Although many studies are devo…

stat.ME2018

Resampling Strategy in Sequential Monte Carlo for Constrained Sampling Problems

Chencheng Cai, Rong Chen, Ming Lin

Sequential Monte Carlo (SMC) methods are a class of Monte Carlo methods that are used to obtain random samples of a high dimensional random variable in a sequential fashion. Many p…

cs.CY2025

PHORECAST: Enabling AI Understanding of Public Health Outreach Across Populations

Rifaa Qadri, Anh Nhat Nhu, Swati Ramnath +6

Understanding how diverse individuals and communities respond to persuasive messaging holds significant potential for advancing personalized and socially aware machine learning. Wh…

cs.CV2023

PAC-NeRF: Physics Augmented Continuum Neural Radiance Fields for Geometry-Agnostic System Identification

Xuan Li, Yi-Ling Qiao, Peter Yichen Chen +4

Existing approaches to system identification (estimating the physical parameters of an object) from videos assume known object geometries. This precludes their applicability in a v…

cs.CV2017

Self-paced Convolutional Neural Network for Computer Aided Detection in Medical Imaging Analysis

Xiang Li, Aoxiao Zhong, Ming Lin +6

Tissue characterization has long been an important component of Computer Aided Diagnosis (CAD) systems for automatic lesion detection and further clinical planning. Motivated by th…

cs.CV2023

Aerial Diffusion: Text Guided Ground-to-Aerial View Translation from a Single Image using Diffusion Models

Divya Kothandaraman, Tianyi Zhou, Ming Lin +1

We present a novel method, Aerial Diffusion, for generating aerial views from a single ground-view image using text guidance. Aerial Diffusion leverages a pretrained text-image dif…

cs.CV2015

Long-short Term Motion Feature for Action Classification and Retrieval

Zhenzhong Lan, Xuanchong Li, Ming Lin +1

We propose a method for representing motion information for video classification and retrieval. We improve upon local descriptor based methods that have been among the most popular…

cs.LG2021

Fine-Grained AutoAugmentation for Multi-Label Classification

Ya Wang, Hesen Chen, Fangyi Zhang +4

Data augmentation is a commonly used approach to improving the generalization of deep learning models. Recent works show that learned data augmentation policies can achieve better…

cs.CV2015

The Best of Both Worlds: Combining Data-independent and Data-driven Approaches for Action Recognition

Zhenzhong Lan, Dezhong Yao, Ming Lin +2

Motivated by the success of data-driven convolutional neural networks (CNNs) in object recognition on static images, researchers are working hard towards developing CNN equivalents…

cs.LG2024

Zero-Shot Neural Architecture Search: Challenges, Solutions, and Opportunities

Guihong Li, Duc Hoang, Kartikeya Bhardwaj +3

Recently, zero-shot (or training-free) Neural Architecture Search (NAS) approaches have been proposed to liberate NAS from the expensive training process. The key idea behind zero-…

eess.IV2022

Entroformer: A Transformer-based Entropy Model for Learned Image Compression

Yichen Qian, Ming Lin, Xiuyu Sun +2

One critical component in lossy deep image compression is the entropy model, which predicts the probability distribution of the quantized latent representation in the encoding and…

cs.CV2022

Differentiable Frequency-based Disentanglement for Aerial Video Action Recognition

Divya Kothandaraman, Ming Lin, Dinesh Manocha

We present a learning algorithm for human activity recognition in videos. Our approach is designed for UAV videos, which are mainly acquired from obliquely placed dynamic cameras t…

cond-mat.mes-hall2022

Unveiling the emergent traits of chiral spin textures in magnetic multilayers

Xiaoye Chen, Ming Lin, Jian Feng Kong +7

Magnetic skyrmions are topologically wound nanoscale textures of spins whose ambient stability and electrical manipulation in multilayer films have led to an explosion of research…

stat.ML2020

Robust Finite Mixture Regression for Heterogeneous Targets

Jian Liang, Kun Chen, Ming Lin +2

Finite Mixture Regression (FMR) refers to the mixture modeling scheme which learns multiple regression models from the training data set. Each of them is in charge of a subset. FMR…

cs.CY2017

Research Opportunities and Visions for Smart and Pervasive Health

Elizabeth Mynatt, Gregory D. Hager, Santosh Kumar +4

Improving the health of the nation's population and increasing the capabilities of the US healthcare system to support diagnosis, treatment, and prevention of disease is a critical…

cs.LG2020

Knapsack Pruning with Inner Distillation

Yonathan Aflalo, Asaf Noy, Ming Lin +2

Neural network pruning reduces the computational cost of an over-parameterized network to improve its efficiency. Popular methods vary from -norm sparsification to Neural A…

cond-mat.stat-mech2026

Mesoscopic MCT theory resolves Giant Non-Gaussian Parameter and Flory's conjecture

Yikun Ren, Feixiang Xu, Ming Lin

Extending Prigogine's ideas to the interior of the system, we generalize mode-coupling theory from a microscopic to a mesoscopic formulation by incorporating the non-equilibrium ei…

cs.CV2024

HawkI: Homography & Mutual Information Guidance for 3D-free Single Image to Aerial View

Divya Kothandaraman, Tianyi Zhou, Ming Lin +1

We present HawkI, for synthesizing aerial-view images from text and an exemplar image, without any additional multi-view or 3D information for finetuning or at inference. HawkI use…

cs.LG2020

WeMix: How to Better Utilize Data Augmentation

Yi Xu, Asaf Noy, Ming Lin +3

Data augmentation is a widely used training trick in deep learning to improve the network generalization ability. Despite many encouraging results, several recent studies did point…

cs.CV2021

Zen-NAS: A Zero-Shot NAS for High-Performance Deep Image Recognition

Ming Lin, Pichao Wang, Zhenhong Sun +5

Accuracy predictor is a key component in Neural Architecture Search (NAS) for ranking architectures. Building a high-quality accuracy predictor usually costs enormous computation.…

cs.CL2025

ClimateChat: Designing Data and Methods for Instruction Tuning LLMs to Answer Climate Change Queries

Zhou Chen, Xiao Wang, Yuanhong Liao +2

As the issue of global climate change becomes increasingly severe, the demand for research in climate science continues to grow. Natural language processing technologies, represent…

cs.CV2024

ViLA: Efficient Video-Language Alignment for Video Question Answering

Xijun Wang, Junbang Liang, Chun-Kai Wang +4

In this work, we propose an efficient Video-Language Alignment (ViLA) network. Our ViLA model addresses both efficient frame sampling and effective cross-modal alignment in a unifi…

cs.LG2025

Adaptive Conformal Guidance for Learning under Uncertainty

Rui Liu, Peng Gao, Yu Shen +2

Learning with guidance has proven effective across a wide range of machine learning systems. Guidance may, for example, come from annotated datasets in supervised learning, pseudo-…

cs.AI2026

DRPO: Efficient Reasoning via Decoupled Reward Policy Optimization

Gang Li, Yan Chen, Ming Lin +1

Recent large reasoning models (LRMs) driven by reinforcement learning algorithms (e.g., GRPO) have achieved remarkable performance on challenging reasoning tasks. However, these mo…

cs.CV2023

ICAR: Image-based Complementary Auto Reasoning

Xijun Wang, Anqi Liang, Junbang Liang +3

Scene-aware Complementary Item Retrieval (CIR) is a challenging task which requires to generate a set of compatible items across domains. Due to the subjectivity, it is difficult t…

cs.SD2017

Effects of virtual acoustics on dynamic auditory distance perception

Atul Rungta, Nicholas Rewkowski, Roberta Klatzky +2

Sound propagation encompasses various acoustic phenomena including reverberation. Current virtual acoustic methods, ranging from parametric filters to physically-accurate solvers,…

cs.CV2023

Making Vision Transformers Efficient from A Token Sparsification View

Shuning Chang, Pichao Wang, Ming Lin +4

The quadratic computational complexity to the number of tokens limits the practical applications of Vision Transformers (ViTs). Several works propose to prune redundant tokens to a…

cs.GR2025

Graphics4Science: Computer Graphics for Scientific Impacts

Peter Yichen Chen, Minghao Guo, Hanspeter Pfister +5

Computer graphics, often associated with films, games, and visual effects, has long been a powerful tool for addressing scientific challenges--from its origins in 3D visualization…

stat.ML2017

Nonconvex One-bit Single-label Multi-label Learning

Shuang Qiu, Tingjin Luo, Jieping Ye +1

We study an extreme scenario in multi-label learning where each training instance is endowed with a single one-bit label out of multiple labels. We formulate this problem as a non-…

cs.CV2022

Ada-NETS: Face Clustering via Adaptive Neighbour Discovery in the Structure Space

Yaohua Wang, Yaobin Zhang, Fangyi Zhang +4

Face clustering has attracted rising research interest recently to take advantage of massive amounts of face images on the web. State-of-the-art performance has been achieved by Gr…

cs.CV2022

KVT: k-NN Attention for Boosting Vision Transformers

Pichao Wang, Xue Wang, Fan Wang +4

Convolutional Neural Networks (CNNs) have dominated computer vision for years, due to its ability in capturing locality and translation invariance. Recently, many vision transforme…

cond-mat.mes-hall2024

Giant third-order nonlinear Hall effect in misfit layer compound (SnS)(NbS)

Shengyao Li, Xueyan Wang, Zherui Yang +11

Nonlinear Hall effect (NLHE) holds immense significance in recognizing the band geometry and its potential applications in current rectification. Recent discoveries have expanded t…

cs.CV2023

SHARE: Single-view Human Adversarial REconstruction

Shreelekha Revankar, Shijia Liao, Yu Shen +3

The accuracy of 3D Human Pose and Shape reconstruction (HPS) from an image is progressively improving. Yet, no known method is robust across all image distortion. To address issues…

cs.CV2022

GiraffeDet: A Heavy-Neck Paradigm for Object Detection

Yiqi Jiang, Zhiyu Tan, Junyan Wang +3

In conventional object detection frameworks, a backbone body inherited from image recognition models extracts deep latent features and then a neck module fuses these latent feature…

cs.AI2024

Search for Efficient Large Language Models

Xuan Shen, Pu Zhao, Yifan Gong +7

Large Language Models (LLMs) have long held sway in the realms of artificial intelligence research. Numerous efficient techniques, including weight pruning, quantization, and disti…

stat.ML2016

A Non-convex One-Pass Framework for Generalized Factorization Machine and Rank-One Matrix Sensing

Ming Lin, Jieping Ye

We develop an efficient alternating framework for learning a generalized version of Factorization Machine (gFM) on steaming data with provable guarantees. When the instances are sa…

cs.CV2025

Financial Models in Generative Art: Black-Scholes-Inspired Concept Blending in Text-to-Image Diffusion

Divya Kothandaraman, Ming Lin, Dinesh Manocha

We introduce a novel approach for concept blending in pretrained text-to-image diffusion models, aiming to generate images at the intersection of multiple text prompts. At each tim…

cs.CV2023

HandyPriors: Physically Consistent Perception of Hand-Object Interactions with Differentiable Priors

Shutong Zhang, Yi-Ling Qiao, Guanglei Zhu +6

Various heuristic objectives for modeling hand-object interaction have been proposed in past work. However, due to the lack of a cohesive framework, these objectives often possess…

cs.LG2026

DisCO: Reinforcing Large Reasoning Models with Discriminative Constrained Optimization

Gang Li, Ming Lin, Tomer Galanti +2

The recent success and openness of DeepSeek-R1 have brought widespread attention to Group Relative Policy Optimization (GRPO) as a reinforcement learning method for large reasoning…

cs.AI2025

MMCD: Multi-Modal Collaborative Decision-Making for Connected Autonomy with Knowledge Distillation

Rui Liu, Zikang Wang, Peng Gao +3

Autonomous systems have advanced significantly, but challenges persist in accident-prone environments where robust decision-making is crucial. A single vehicle's limited sensor ran…

quant-ph2022

Differentiable Analog Quantum Computing for Optimization and Control

Jiaqi Leng, Yuxiang Peng, Yi-Ling Qiao +2

We formulate the first differentiable analog quantum computing framework with a specific parameterization design at the analog signal (pulse) level to better exploit near-term quan…

cs.RO2025

CAML: Collaborative Auxiliary Modality Learning for Multi-Agent Systems

Rui Liu, Yu Shen, Peng Gao +2

Multi-modal learning has emerged as a key technique for improving performance across domains such as autonomous driving, robotics, and reasoning. However, in certain scenarios, par…

cs.CV2022

Robust Graph Structure Learning via Multiple Statistical Tests

Yaohua Wang, FangYi Zhang, Ming Lin +3

Graph structure learning aims to learn connectivity in a graph from data. It is particularly important for many computer vision related tasks since no explicit graph structure is a…

cs.CV2023

Learning the Relation between Similarity Loss and Clustering Loss in Self-Supervised Learning

Jidong Ge, Yuxiang Liu, Jie Gui +5

Self-supervised learning enables networks to learn discriminative features from massive data itself. Most state-of-the-art methods maximize the similarity between two augmentations…

cs.LG2026

Active Asymmetric Multi-Agent Multimodal Learning under Uncertainty

Rui Liu, Pratap Tokekar, Ming Lin

Multi-agent systems are increasingly equipped with heterogeneous multimodal sensors, enabling richer perception but introducing modality-specific and agent-dependent uncertainty. E…

cs.LG2025

Agile-Quant: Activation-Guided Quantization for Faster Inference of LLMs on the Edge

Xuan Shen, Peiyan Dong, Lei Lu +5

Large Language Models (LLMs) stand out for their impressive performance in intricate language modeling tasks. However, their demanding computational and memory needs pose obstacles…

cs.RO2022

WGICP: Differentiable Weighted GICP-Based Lidar Odometry

Sanghyun Son, Jing Liang, Ming Lin +1

We present a novel differentiable weighted generalized iterative closest point (WGICP) method applicable to general 3D point cloud data, including that from Lidar. Our method build…

cs.LG2026

Beyond Test-Time Memory: State-Space Optimal Control for LLM Reasoning

Peihao Wang, Shan Yang, Xijun Wang +8

Associative memory has long underpinned the design of sequential models. Beyond recall, humans reason by projecting future states and selecting goal-directed actions, a capability…

cs.CV2022

FAR: Fourier Aerial Video Recognition

Divya Kothandaraman, Tianrui Guan, Xijun Wang +3

We present an algorithm, Fourier Activity Recognition (FAR), for UAV video activity recognition. Our formulation uses a novel Fourier object disentanglement method to innately sepa…

cs.CV2015

Beyond Gaussian Pyramid: Multi-skip Feature Stacking for Action Recognition

Zhenzhong Lan, Ming Lin, Xuanchong Li +2

Most state-of-the-art action feature extractors involve differential operators, which act as highpass filters and tend to attenuate low frequency action information. This attenuati…

cs.IR2016

Strategies for Searching Video Content with Text Queries or Video Examples

Shoou-I Yu, Yi Yang, Zhongwen Xu +13

The large number of user-generated videos uploaded on to the Internet everyday has led to many commercial video search engines, which mainly rely on text metadata for search. Howev…

cs.CV2025

HART: Human Aligned Reconstruction Transformer

Xiyi Chen, Shaofei Wang, Marko Mihajlovic +3

We introduce HART, a unified framework for sparse-view human reconstruction. Given a small set of uncalibrated RGB images of a person as input, it outputs a watertight clothed mesh…

cs.CV2015

Handcrafted Local Features are Convolutional Neural Networks

Zhenzhong Lan, Shoou-I Yu, Ming Lin +2

Image and video classification research has made great progress through the development of handcrafted local features and learning based features. These two architectures were prop…

cs.AI2025

CharCom: Composable Identity Control for Multi-Character Story Illustration

Zhongsheng Wang, Ming Lin, Zhedong Lin +3

Ensuring character identity consistency across varying prompts remains a fundamental limitation in diffusion-based text-to-image generation. We propose CharCom, a modular and param…

cs.LG2025

Time-Aware World Model for Adaptive Prediction and Control

Anh N. Nhu, Sanghyun Son, Ming Lin

In this work, we introduce the Time-Aware World Model (TAWM), a model-based approach that explicitly incorporates temporal dynamics. By conditioning on the time-step size, Δt, and…

cs.CV2022

MAE-DET: Revisiting Maximum Entropy Principle in Zero-Shot NAS for Efficient Object Detection

Zhenhong Sun, Ming Lin, Xiuyu Sun +3

In object detection, the detection backbone consumes more than half of the overall inference cost. Recent researches attempt to reduce this cost by optimizing the backbone architec…

cs.LG2025

Merino: Entropy-driven Design for Generative Language Models on IoT Devices

Youpeng Zhao, Ming Lin, Huadong Tang +2

Generative Large Language Models (LLMs) stand as a revolutionary advancement in the modern era of artificial intelligence (AI). However, scaling down LLMs for resource-constrained…

cs.RO2020

Enhanced Transfer Learning for Autonomous Driving with Systematic Accident Simulation

Shivam Akhauri, Laura Zheng, Ming Lin

Simulation data can be utilized to extend real-world driving data in order to cover edge cases, such as vehicle accidents. The importance of handling edge cases can be observed in…

stat.ME2013

Lookahead Strategies for Sequential Monte Carlo

Ming Lin, Rong Chen, Jun S. Liu

Based on the principles of importance sampling and resampling, sequential Monte Carlo (SMC) encompasses a large set of powerful techniques dealing with complex stochastic dynamic s…

cond-mat.stat-mech2025

Non-Equilibrium Thermodynamics Framework to Address the Glass Transition

Yikun Ren, Feixiang Xu, Ming Lin

When the center of fluctuations, i.e., the nonequilibrium eigenphase, undergoes transformation, there emerge critical parameters that demonstrate insensitivity to fluctuation pertu…