Publications (102)
RTMobile: Beyond Real-Time Mobile Acceleration of RNNs for Speech Recognition
Peiyan Dong, Siyue Wang, Wei Niu +8
Recurrent neural networks (RNNs) based automatic speech recognition has nowadays become prevalent on mobile devices such as smart phones. However, previous RNN compression techniqu…
MEST: Accurate and Fast Memory-Economic Sparse Training Framework on the Edge
Geng Yuan, Xiaolong Ma, Wei Niu +13
Recently, a new trend of exploring sparsity for accelerating neural network training has emerged, embracing the paradigm of training on the edge. This paper proposes a novel Memory…
Structured Agent Distillation for Large Language Model
Jun Liu, Zhenglun Kong, Peiyan Dong +10
Large language models (LLMs) exhibit strong capabilities as decision-making agents by interleaving reasoning and actions, as seen in ReAct-style frameworks. Yet, their practical de…
Beta-CROWN: Efficient Bound Propagation with Per-neuron Split Constraints for Complete and Incomplete Neural Network Robustness Verification
Shiqi Wang, Huan Zhang, Kaidi Xu +4
Bound propagation based incomplete neural network verifiers such as CROWN are very efficient and can significantly accelerate branch-and-bound (BaB) based complete verification of…
Progressive Weight Pruning of Deep Neural Networks using ADMM
Shaokai Ye, Tianyun Zhang, Kaiqi Zhang +10
Deep neural networks (DNNs) although achieving human-level performance in many domains, have very large model size that hinders their broader applications on edge computing devices…
LUTMUL: Exceed Conventional FPGA Roofline Limit by LUT-based Efficient Multiplication for Neural Network Inference
Yanyue Xie, Zhengang Li, Dana Diaconu +3
For FPGA-based neural network accelerators, digital signal processing (DSP) blocks have traditionally been the cornerstone for handling multiplications. This paper introduces LUTMU…
Dirty Road Can Attack: Security of Deep Learning based Automated Lane Centering under Physical-World Attack
Takami Sato, Junjie Shen, Ningfei Wang +3
Automated Lane Centering (ALC) systems are convenient and widely deployed today, but also highly security and safety critical. In this work, we are the first to systematically stud…
Achieving Real-Time LiDAR 3D Object Detection on a Mobile Device
Pu Zhao, Wei Niu, Geng Yuan +7
3D object detection is an important task, especially in the autonomous driving application domain. However, it is challenging to support the real-time performance with the limited…
BLK-REW: A Unified Block-based DNN Pruning Framework using Reweighted Regularization Method
Xiaolong Ma, Zhengang Li, Yifan Gong +8
Accelerating DNN execution on various resource-limited computing platforms has been a long-standing problem. Prior works utilize l1-based group lasso or dynamic regularization such…
VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting
Juyi Lin, Amir Taherin, Arash Akbari +11
Recent large-scale Vision Language Action (VLA) models have shown superior performance in robotic manipulation tasks guided by natural language. However, current VLA models suffer…
ILMPQ : An Intra-Layer Multi-Precision Deep Neural Network Quantization framework for FPGA
Sung-En Chang, Yanyu Li, Mengshu Sun +2
This work targets the commonly used FPGA (field-programmable gate array) devices as the hardware platform for DNN edge computing. We focus on DNN quantization as the main model com…
NPAS: A Compiler-aware Framework of Unified Network Pruning and Architecture Search for Beyond Real-Time Mobile Acceleration
Zhengang Li, Geng Yuan, Wei Niu +13
With the increasing demand to efficiently deploy DNNs on mobile edge devices, it becomes much more important to reduce unnecessary computation and increase the execution speed. Pri…
StructADMM: A Systematic, High-Efficiency Framework of Structured Weight Pruning for DNNs
Tianyun Zhang, Shaokai Ye, Kaiqi Zhang +8
Weight pruning methods of DNNs have been demonstrated to achieve a good model pruning rate without loss of accuracy, thereby alleviating the significant computation/storage require…
FAIVConf: Face enhancement for AI-based Video Conference with Low Bit-rate
Zhengang Li, Sheng Lin, Shan Liu +4
Recently, high-quality video conferencing with fewer transmission bits has become a very hot and challenging problem. We propose FAIVConf, a specially designed video compression fr…
Location-free Human Pose Estimation
Xixia Xu, Yingguo Gao, Ke Yan +2
Human pose estimation (HPE) usually requires large-scale training data to reach high performance. However, it is rather time-consuming to collect high-quality and fine-grained anno…
Machine-learning-assisted electron-spin readout of nitrogen-vacancy center in diamond
Peng Qian, Xue Lin, Feifei Zhou +5
Machine learning is a powerful tool in finding hidden data patterns for quantum information processing. Here, we introduce this method into the optical readout of electron-spin sta…
Can Adversarial Examples Be Parsed to Reveal Victim Model Information?
Yuguang Yao, Jiancheng Liu, Yifan Gong +4
Numerous adversarial attack methods have been developed to generate imperceptible image perturbations that can cause erroneous predictions of state-of-the-art machine learning (ML)…
Achieving on-Mobile Real-Time Super-Resolution with Neural Architecture and Pruning Search
Zheng Zhan, Yifan Gong, Pu Zhao +9
Though recent years have witnessed remarkable progress in single image super-resolution (SISR) tasks with the prosperous development of deep neural networks (DNNs), the deep learni…
Automatic Mapping of the Best-Suited DNN Pruning Schemes for Real-Time Mobile Acceleration
Yifan Gong, Geng Yuan, Zheng Zhan +9
Weight pruning is an effective model compression technique to tackle the challenges of achieving real-time deep neural network (DNN) inference on mobile devices. However, prior pru…
High-Robustness, Low-Transferability Fingerprinting of Neural Networks
Siyue Wang, Xiao Wang, Pin-Yu Chen +2
This paper proposes Characteristic Examples for effectively fingerprinting deep neural networks, featuring high-robustness to the base model against model pruning as well as low-tr…
Fast and Complete: Enabling Complete Neural Network Verification with Rapid and Massively Parallel Incomplete Verifiers
Kaidi Xu, Huan Zhang, Shiqi Wang +4
Formal verification of neural networks (NNs) is a challenging and important problem. Existing efficient complete solvers typically require the branch-and-bound (BaB) process, which…
AdvMS: A Multi-source Multi-cost Defense Against Adversarial Attacks
Xiao Wang, Siyue Wang, Pin-Yu Chen +2
Designing effective defense against adversarial attacks is a crucial topic as deep neural networks have been proliferated rapidly in many security-critical domains such as malware…
Structured Adversarial Attack: Towards General Implementation and Better Interpretability
Kaidi Xu, Sijia Liu, Pu Zhao +6
When generating adversarial examples to attack deep neural networks (DNNs), Lp norm of the added perturbation is usually used to measure the similarity between original image and a…
Noise-Resilient Quantum Metrology with Quantum Computing
Xiangyu Wang, Chenrong Liu, Xue Lin +8
Quantum computing has made remarkable strides in recent years, as demonstrated by quantum supremacy experiments and the realization of high-fidelity, fault-tolerant gates. However,…
Adversarial Robustness vs Model Compression, or Both?
Shaokai Ye, Kaidi Xu, Sijia Liu +6
It is well known that deep neural networks (DNNs) are vulnerable to adversarial attacks, which are implemented by adding crafted perturbations onto benign examples. Min-max robust…
Brain Tumor Classification on MRI in Light of Molecular Markers
Jun Liu, Geng Yuan, Weihao Zeng +6
In research findings, co-deletion of the 1p/19q gene is associated with clinical outcomes in low-grade gliomas. The ability to predict 1p19q status is critical for treatment planni…
Online optimization for optical readout of a single electron spin in diamond
Xue Lin, Jingwei Fan, Runchuan Ye +4
The nitrogen-vacancy (NV) center in diamond has been developed as a promising platform for quantum sensing, especially for magnetic field measurements in the nano-tesla range with…
Noise prediction and reduction of single electron spin by deep-learning-enhanced feedforward control
Nanyang Xu, Feifei Zhou, Xiangyu Ye +7
Noise-induced control imperfection is an important problem in applications of diamond-based nano-scale sensing, where measurement-based strategies are generally utilized to correct…
Prompt-based Adaptation in Large-scale Vision Models: A Survey
Xi Xiao, Yunbei Zhang, Lin Zhao +12
In computer vision, Visual Prompting (VP) and Visual Prompt Tuning (VPT) have recently emerged as lightweight and effective alternatives to full fine-tuning for adapting large-scal…
Bridging Mode Connectivity in Loss Landscapes and Adversarial Robustness
Pu Zhao, Pin-Yu Chen, Payel Das +2
Mode connectivity provides novel geometric insights on analyzing loss landscapes and enables building high-accuracy pathways between well-trained neural networks. In this work, we…
SuperFlow: A Fully-Customized RTL-to-GDS Design Automation Flow for Adiabatic Quantum-Flux-Parametron Superconducting Circuits
Yanyue Xie, Peiyan Dong, Geng Yuan +10
Superconducting circuits, like Adiabatic Quantum-Flux-Parametron (AQFP), offer exceptional energy efficiency but face challenges in physical design due to sophisticated spacing and…
Unsafer in Many Turns: Benchmarking and Defending Multi-Turn Safety Risks in Tool-Using Agents
Xu Li, Simon Yu, Minzhou Pan +5
LLM-based agents are becoming increasingly capable, yet their safety lags behind. This creates a gap between what agents can do and should do. This gap widens as agents engage in m…
Taming Diffusion for Dataset Distillation with High Representativeness
Lin Zhao, Yushu Wu, Xinru Jiang +5
Recent deep learning models demand larger datasets, driving the need for dataset distillation to create compact, cost-efficient datasets while maintaining performance. Due to the p…
RoRA: Efficient Fine-Tuning of LLM with Reliability Optimization for Rank Adaptation
Jun Liu, Zhenglun Kong, Peiyan Dong +10
Fine-tuning helps large language models (LLM) recover degraded information and enhance task performance. Although Low-Rank Adaptation (LoRA) is widely used and effective for fine-t…
Security of Deep Learning based Lane Keeping System under Physical-World Adversarial Attack
Takami Sato, Junjie Shen, Ningfei Wang +3
Lane-Keeping Assistance System (LKAS) is convenient and widely available today, but also extremely security and safety critical. In this work, we design and implement the first sys…
MoE-Pruner: Pruning Mixture-of-Experts Large Language Model using the Hints from Its Router
Yanyue Xie, Zhi Zhang, Ding Zhou +6
Mixture-of-Experts (MoE) architectures face challenges such as high memory consumption and redundancy in experts. Pruning MoE can reduce network weights while maintaining model per…
Pruning then Reweighting: Towards Data-Efficient Training of Diffusion Models
Yize Li, Yihua Zhang, Sijia Liu +1
Despite the remarkable generation capabilities of Diffusion Models (DMs), conducting training and inference remains computationally expensive. Previous works have been devoted to a…
Finding needles in a haystack: A Black-Box Approach to Invisible Watermark Detection
Minzhou Pan, Zhenting Wang, Xin Dong +3
In this paper, we propose WaterMark Detection (WMD), the first invisible watermark detection method under a black-box and annotation-free setting. WMD is capable of detecting arbit…
Towards an Efficient and General Framework of Robust Training for Graph Neural Networks
Kaidi Xu, Sijia Liu, Pin-Yu Chen +4
Graph Neural Networks (GNNs) have made significant advances on several fundamental inference tasks. As a result, there is a surge of interest in using these models for making poten…
RCR-Router: Efficient Role-Aware Context Routing for Multi-Agent LLM Systems with Structured Memory
Jun Liu, Zhenglun Kong, Changdi Yang +12
Multi-agent large language model (LLM) systems have shown strong potential in complex reasoning and collaborative decision-making tasks. However, most existing coordination schemes…
Fault Sneaking Attack: a Stealthy Framework for Misleading Deep Neural Networks
Pu Zhao, Siyue Wang, Cheng Gongye +3
Despite the great achievements of deep neural networks (DNNs), the vulnerability of state-of-the-art DNNs raises security concerns of DNNs in many application domains requiring hig…
Alleviating Human-level Shift : A Robust Domain Adaptation Method for Multi-person Pose Estimation
Xixia Xu, Qi Zou, Xue Lin
Human pose estimation has been widely studied with much focus on supervised learning requiring sufficient annotations. However, in real applications, a pretrained pose estimation m…
ADMM-NN: An Algorithm-Hardware Co-Design Framework of DNNs Using Alternating Direction Method of Multipliers
Ao Ren, Tianyun Zhang, Shaokai Ye +5
To facilitate efficient embedded and hardware implementations of deep neural networks (DNNs), two important categories of DNN model compression techniques: weight pruning and weigh…
Feature Distillation: DNN-Oriented JPEG Compression Against Adversarial Examples
Zihao Liu, Qi Liu, Tao Liu +4
Image compression-based approaches for defending against the adversarial-example attacks, which threaten the safety use of deep neural networks (DNN), have been investigated recent…
A Unified Framework of DNN Weight Pruning and Weight Clustering/Quantization Using ADMM
Shaokai Ye, Tianyun Zhang, Kaiqi Zhang +6
Many model compression techniques of Deep Neural Networks (DNNs) have been investigated, including weight pruning, weight clustering and quantization, etc. Weight pruning leverages…
PatDNN: Achieving Real-Time DNN Execution on Mobile Devices with Pattern-based Weight Pruning
Wei Niu, Xiaolong Ma, Sheng Lin +5
With the emergence of a spectrum of high-end mobile devices, many applications that formerly required desktop-level computation capability are being transferred to these devices. H…
Topology Attack and Defense for Graph Neural Networks: An Optimization Perspective
Kaidi Xu, Hongge Chen, Sijia Liu +4
Graph neural networks (GNNs) which apply the deep neural networks to graph data have achieved significant performance for the task of semi-supervised node classification. However,…
Towards Ultra-High Performance and Energy Efficiency of Deep Learning Systems: An Algorithm-Hardware Co-Optimization Framework
Yanzhi Wang, Caiwen Ding, Zhe Li +8
Hardware accelerations of deep learning systems have been extensively investigated in industry and academia. The aim of this paper is to achieve ultra-high energy efficiency and pe…
Reweighted Proximal Pruning for Large-Scale Language Representation
Fu-Ming Guo, Sijia Liu, Finlay S. Mungall +2
Recently, pre-trained language representation flourishes as the mainstay of the natural language understanding community, e.g., BERT. These pre-trained language representations can…
Pruning Foundation Models for High Accuracy without Retraining
Pu Zhao, Fei Sun, Xuan Shen +4
Despite the superior performance, it is challenging to deploy foundation models or large language models (LLMs) due to their massive parameters and computations. While pruning is a…
Learning to Generate Image Source-Agnostic Universal Adversarial Perturbations
Pu Zhao, Parikshit Ram, Songtao Lu +4
Adversarial perturbations are critical for certifying the robustness of deep learning models. A universal adversarial perturbation (UAP) can simultaneously attack multiple images,…
CirCNN: Accelerating and Compressing Deep Neural Networks Using Block-CirculantWeight Matrices
Caiwen Ding, Siyu Liao, Yanzhi Wang +13
Large-scale deep neural networks (DNNs) are both compute and memory intensive. As the size of DNNs continues to grow, it is critical to improve the energy efficiency and performanc…
Toward Adaptive Large Language Models Structured Pruning via Hybrid-grained Weight Importance Assessment
Jun Liu, Zhenglun Kong, Pu Zhao +9
Structured pruning for large language models (LLMs) has garnered significant academic interest due to its ability to efficiently compress and accelerate LLMs by eliminating redunda…
Block Switching: A Stochastic Approach for Deep Learning Security
Xiao Wang, Siyue Wang, Pin-Yu Chen +2
Recent study of adversarial attacks has revealed the vulnerability of modern deep learning models. That is, subtly crafted perturbations of the input can make a trained network wit…
Defensive Dropout for Hardening Deep Neural Networks under Adversarial Attacks
Siyue Wang, Xiao Wang, Pu Zhao +4
Deep neural networks (DNNs) are known vulnerable to adversarial attacks. That is, adversarial examples, obtained by adding delicately crafted distortions onto original legal inputs…
Zeroth-Order Hybrid Gradient Descent: Towards A Principled Black-Box Optimization Framework
Pranay Sharma, Kaidi Xu, Sijia Liu +3
In this work, we focus on the study of stochastic zeroth-order (ZO) optimization which does not require first-order gradient information and uses only function evaluations. The pro…
Efficient Multi-Prize Lottery Tickets: Enhanced Accuracy, Training, and Inference Speed
Hao Cheng, Pu Zhao, Yize Li +4
Recently, Diffenderfer and Kailkhura proposed a new paradigm for learning compact yet highly accurate binary neural networks simply by pruning and quantizing randomly weighted full…
A Privacy-Preserving-Oriented DNN Pruning and Mobile Acceleration Framework
Yifan Gong, Zheng Zhan, Zhengang Li +8
Weight pruning of deep neural networks (DNNs) has been proposed to satisfy the limited storage and computing capability of mobile edge devices. However, previous pruning methods ma…
GRIM: A General, Real-Time Deep Learning Inference Framework for Mobile Devices based on Fine-Grained Structured Weight Sparsity
Wei Niu, Zhengang Li, Xiaolong Ma +6
It is appealing but challenging to achieve real-time deep neural network (DNN) inference on mobile devices because even the powerful modern mobile devices are considered as ``resou…
Mix and Match: A Novel FPGA-Centric Deep Neural Network Quantization Framework
Sung-En Chang, Yanyu Li, Mengshu Sun +5
Deep Neural Networks (DNNs) have achieved extraordinary performance in various application domains. To support diverse DNN models, efficient implementations of DNN inference on edg…
MSP: An FPGA-Specific Mixed-Scheme, Multi-Precision Deep Neural Network Quantization Framework
Sung-En Chang, Yanyu Li, Mengshu Sun +4
With the tremendous success of deep learning, there exists imminent need to deploy deep learning models onto edge devices. To tackle the limited computing and storage resources in…
Search for Efficient Large Language Models
Xuan Shen, Pu Zhao, Yifan Gong +7
Large Language Models (LLMs) have long held sway in the realms of artificial intelligence research. Numerous efficient techniques, including weight pruning, quantization, and disti…
ActQuant: Sub-4-bit Action-Guided Quantization for Vision-Language-Action Models
Arash Akbari, Arman Akbari, Masih Eskandar +11
Vision-Language-Action (VLA) models exhibit remarkable action generation for embodied intelligence, but their heavy compute make deployment on edge platforms impractical. Aggressiv…
Pursing the Sparse Limitation of Spiking Deep Learning Structures
Hao Cheng, Jiahang Cao, Erjia Xiao +7
Spiking Neural Networks (SNNs), a novel brain-inspired algorithm, are garnering increased attention for their superior computation and energy efficiency over traditional artificial…
Interpreting Adversarial Examples by Activation Promotion and Suppression
Kaidi Xu, Sijia Liu, Gaoyuan Zhang +5
It is widely known that convolutional neural networks (CNNs) are vulnerable to adversarial examples: images with imperceptible perturbations crafted to fool classifiers. However, i…
Achieving Real-Time Object Detection on MobileDevices with Neural Pruning Search
Pu Zhao, Wei Niu, Geng Yuan +4
Object detection plays an important role in self-driving cars for security development. However, mobile systems on self-driving cars with limited computation resources lead to diff…
Pruning-as-Search: Efficient Neural Architecture Search via Channel Pruning and Structural Reparameterization
Yanyu Li, Pu Zhao, Geng Yuan +3
Neural architecture search (NAS) and network pruning are widely studied efficient AI techniques, but not yet perfect. NAS performs exhaustive candidate architecture search, incurri…
Auto-ViT-Acc: An FPGA-Aware Automatic Acceleration Framework for Vision Transformer with Mixed-Scheme Quantization
Zhengang Li, Mengshu Sun, Alec Lu +9
Vision transformers (ViTs) are emerging with significantly improved accuracy in computer vision tasks. However, their complex architecture and enormous computation/storage demand i…
Detection and Recovery Against Deep Neural Network Fault Injection Attacks Based on Contrastive Learning
Chenan Wang, Pu Zhao, Siyue Wang +1
Deep Neural Network (DNN) models when implemented on executing devices as the inference engines are susceptible to Fault Injection Attacks (FIAs) that manipulate model parameters t…
Reverse Engineering of Imperceptible Adversarial Image Perturbations
Yifan Gong, Yuguang Yao, Yize Li +4
It has been well recognized that neural network based image classifiers are easily fooled by images with tiny perturbations crafted by an adversary. There has been a vast volume of…
Towards Real-Time DNN Inference on Mobile Platforms with Model Pruning and Compiler Optimization
Wei Niu, Pu Zhao, Zheng Zhan +3
High-end mobile platforms rapidly serve as primary computing devices for a wide range of Deep Neural Network (DNN) applications. However, the constrained computation and storage re…
Mixture of Robust Experts (MoRE):A Robust Denoising Method towards multiple perturbations
Hao Cheng, Kaidi Xu, Chenan Wang +3
To tackle the susceptibility of deep neural networks to adversarial examples, the adversarial training has been proposed which provides a notion of security through an inner maximi…
Progressive DNN Compression: A Key to Achieve Ultra-High Weight Pruning and Quantization Rates using ADMM
Shaokai Ye, Xiaoyu Feng, Tianyun Zhang +11
Weight pruning and weight quantization are two important categories of DNN model compression. Prior work on these techniques are mainly based on heuristics. A recent work developed…
Token Reduction Should Go Beyond Efficiency in Generative Models -- From Vision, Language to Multimodality
Zhenglun Kong, Yize Li, Fanhu Zeng +7
In Transformer architectures, tokens\textemdash discrete units derived from raw data\textemdash are formed by segmenting inputs into fixed-length chunks. Each token is then mapped…
PCONV: The Missing but Desirable Sparsity in DNN Weight Pruning for Real-time Execution on Mobile Devices
Xiaolong Ma, Fu-Ming Guo, Wei Niu +5
Model compression techniques on Deep Neural Network (DNN) have been widely acknowledged as an effective way to achieve acceleration on a variety of platforms, and DNN weight prunin…
Protecting Neural Networks with Hierarchical Random Switching: Towards Better Robustness-Accuracy Trade-off for Stochastic Defenses
Xiao Wang, Siyue Wang, Pin-Yu Chen +4
Despite achieving remarkable success in various domains, recent studies have uncovered the vulnerability of deep neural networks to adversarial perturbations, creating concerns on…
Automatic Perturbation Analysis for Scalable Certified Robustness and Beyond
Kaidi Xu, Zhouxing Shi, Huan Zhang +6
Linear relaxation based perturbation analysis (LiRPA) for neural networks, which computes provable linear bounds of output neurons given a certain amount of input perturbation, has…
Non-Structured DNN Weight Pruning -- Is It Beneficial in Any Platform?
Xiaolong Ma, Sheng Lin, Shaokai Ye +10
Large deep neural network (DNN) models pose the key challenge to energy efficiency due to the significantly higher energy consumption of off-chip DRAM accesses than arithmetic or S…
On the Design of Black-box Adversarial Examples by Leveraging Gradient-free Optimization and Operator Splitting Method
Pu Zhao, Sijia Liu, Pin-Yu Chen +4
Robust machine learning is currently one of the most prominent topics which could potentially help shaping a future of advanced AI platforms that not only perform well in average c…
Multi-Person Pose Estimation with Enhanced Feature Aggregation and Selection
Xixia Xu, Qi Zou, Xue Lin
We propose a novel Enhanced Feature Aggregation and Selection network (EFASNet) for multi-person 2D human pose estimation. Due to enhanced feature representation, our method can we…
RT3D: Achieving Real-Time Execution of 3D Convolutional Neural Networks on Mobile Devices
Wei Niu, Mengshu Sun, Zhengang Li +7
Mobile devices are becoming an important carrier for deep learning tasks, as they are being equipped with powerful, high-end mobile CPUs and GPUs. However, it is still a challengin…
ZO-AdaMM: Zeroth-Order Adaptive Momentum Method for Black-Box Optimization
Xiangyi Chen, Sijia Liu, Kaidi Xu +4
The adaptive momentum method (AdaMM), which uses past gradients to update descent directions and learning rates simultaneously, has become one of the most popular first-order optim…
A Survey of Advancing Audio Super-Resolution and Bandwidth Extension from Discriminative to Generative Models
Ningyuan Yang, Yize Li, Diego A. Cuji +4
Audio super-resolution (SR), also referred to as bandwidth extension (BWE), aims to reconstruct high-fidelity signals from low-resolution (LR) or band-limited (BL) observations, an…
RMSMP: A Novel Deep Neural Network Quantization Framework with Row-wise Mixed Schemes and Multiple Precisions
Sung-En Chang, Yanyu Li, Mengshu Sun +4
This work proposes a novel Deep Neural Network (DNN) quantization framework, namely RMSMP, with a Row-wise Mixed-Scheme and Multi-Precision approach. Specifically, this is the firs…
HeatViT: Hardware-Efficient Adaptive Token Pruning for Vision Transformers
Peiyan Dong, Mengshu Sun, Alec Lu +8
While vision transformers (ViTs) have continuously achieved new milestones in the field of computer vision, their sophisticated network architectures with high computation and memo…
Rethinking Token Reduction for State Space Models
Zheng Zhan, Yushu Wu, Zhenglun Kong +6
Recent advancements in State Space Models (SSMs) have attracted significant interest, particularly in models optimized for parallel training and handling long-range dependencies. A…
Quasar-ViT: Hardware-Oriented Quantization-Aware Architecture Search for Vision Transformers
Zhengang Li, Alec Lu, Yanyue Xie +9
Vision transformers (ViTs) have demonstrated their superior accuracy for computer vision tasks compared to convolutional neural networks (CNNs). However, ViT models are often compu…
HybridFlow: Infusing Continuity into Masked Codebook for Extreme Low-Bitrate Image Compression
Lei Lu, Yanyue Xie, Wei Jiang +3
This paper investigates the challenging problem of learned image compression (LIC) with extreme low bitrates. Previous LIC methods based on transmitting quantized continuous featur…
TSLA: A Task-Specific Learning Adaptation for Semantic Segmentation on Autonomous Vehicles Platform
Jun Liu, Zhenglun Kong, Pu Zhao +9
Autonomous driving platforms encounter diverse driving scenarios, each with varying hardware resources and precision requirements. Given the computational limitations of embedded d…
JIGMARK: A Black-Box Approach for Enhancing Image Watermarks against Diffusion Model Edits
Minzhou Pan, Yi Zeng, Xue Lin +4
In this study, we investigate the vulnerability of image watermarks to diffusion-model-based image editing, a challenge exacerbated by the computational cost of accessing gradient…
ASSET: Robust Backdoor Data Detection Across a Multiplicity of Deep Learning Paradigms
Minzhou Pan, Yi Zeng, Lingjuan Lyu +2
Backdoor data detection is traditionally studied in an end-to-end supervised learning (SL) setting. However, recent years have seen the proliferating adoption of self-supervised le…
Less is More: Data Pruning for Faster Adversarial Training
Yize Li, Pu Zhao, Xue Lin +2
Deep neural networks (DNNs) are sensitive to adversarial examples, resulting in fragile and unreliable performance in the real world. Although adversarial training (AT) is currentl…
Prediction-Based Fast Thermoelectric Generator Reconfiguration for Energy Harvesting from Vehicle Radiators
Hanchen Yang, Feiyang Kang, Caiwen Ding +7
Thermoelectric generation (TEG) has increasingly drawn attention for being environmentally friendly. A few researches have focused on improving TEG efficiency at the system level o…
E-RNN: Design Optimization for Efficient Recurrent Neural Networks in FPGAs
Zhe Li, Caiwen Ding, Siyue Wang +8
Recurrent Neural Networks (RNNs) are becoming increasingly important for time series-related applications which require efficient and real-time implementations. The two major types…
On the Universal Approximation Property and Equivalence of Stochastic Computing-based Neural Networks and Binary Neural Networks
Yanzhi Wang, Zheng Zhan, Jiayu Li +6
Large-scale deep neural networks are both memory intensive and computation-intensive, thereby posing stringent requirements on the computing platforms. Hardware accelerations of de…
Delta-operator based consensus analysis of multi-agent networks with link failures
Xue Lin, Yuanshi Zheng, Long Wang
In this paper, a discrete-time multi-agent system is presented which is formulated in terms of the delta operator. The proposed multi-agent system can unify discrete-time and conti…
Towards Query-Efficient Black-Box Adversary with Zeroth-Order Natural Gradient Descent
Pu Zhao, Pin-Yu Chen, Siyue Wang +1
Despite the great achievements of the modern deep neural networks (DNNs), the vulnerability/robustness of state-of-the-art DNNs raises security concerns in many application domains…
An ADMM-Based Universal Framework for Adversarial Attacks on Deep Neural Networks
Pu Zhao, Sijia Liu, Yanzhi Wang +1
Deep neural networks (DNNs) are known vulnerable to adversarial attacks. That is, adversarial examples, obtained by adding delicately crafted distortions onto original legal inputs…
When Should a Robot Think? Resource-Aware Reasoning via Reinforcement Learning for Embodied Robotic Decision-Making
Jun Liu, Pu Zhao, Zhenglun Kong +12
Embodied robotic systems increasingly rely on large language model (LLM)-based agents to support high-level reasoning, planning, and decision-making during interactions with the en…
Adversarial T-shirt! Evading Person Detectors in A Physical World
Kaidi Xu, Gaoyuan Zhang, Sijia Liu +6
It is known that deep neural networks (DNNs) are vulnerable to adversarial attacks. The so-called physical adversarial examples deceive DNN-based decisionmakers by attaching advers…