papers

Publications (225)

cs.CR2025

Your Trust, Your Terms: A General Paradigm for Near-Instant Cross-Chain Transfer

Di Wu, Jingyu Liu, Xuechao Wang +4

Cross-chain transactions today remain slow, costly, and fragmented. Existing custodial exchanges expose users to counterparty and centralization risks, while non-custodial liquidit…

cs.DC2026

A Few GPUs, A Whole Lotta Scale: Faithful LLM Training Emulation with PrismLLM

Shaoke Xi, ChonLam Lao, Boyi Jia +11

Large language model (LLM) training today runs on clusters spanning thousands of GPUs. While this scale enables rapid model advances, developing, debugging, and performance-tuning…

cs.SD2022

Privacy-Utility Balanced Voice De-Identification Using Adversarial Examples

Meng Chen, Li Lu, Jiadi Yu +4

Faced with the threat of identity leakage during voice data publishing, users are engaged in a privacy-utility dilemma when enjoying convenient voice services. Existing studies emp…

math.OC2025

An Algebraically Converging Stochastic Gradient Descent Algorithm for Global Optimization

Björn Engquist, Kui Ren, Yunan Yang

We propose a new gradient descent algorithm with added stochastic terms for finding the global optimizers of nonconvex optimization problems. A key component in the algorithm is th…

cs.LG2026

Channel-Level Semantic Perturbations: Unlearnable Examples for Diverse Training Paradigms

Bo Wang, Jia Ni, Mengnan Zhao +2

The unauthorized use of personal data in model training has emerged as a growing privacy threat. Unlearnable examples (UEs) address this issue by embedding imperceptible perturbati…

cs.CR2024

Ambush from All Sides: Understanding Security Threats in Open-Source Software CI/CD Pipelines

Ziyue Pan, Wenbo Shen, Xingkai Wang +6

The continuous integration and continuous deployment (CI/CD) pipelines are widely adopted on Internet hosting platforms, such as GitHub. With the popularity, the CI/CD pipeline fac…

cs.CV2025

Robust Watermarks Leak: Channel-Aware Feature Extraction Enables Adversarial Watermark Manipulation

Zhongjie Ba, Yitao Zhang, Peng Cheng +4

Watermarking plays a key role in the provenance and detection of AI-generated content. While existing methods prioritize robustness against real-world distortions (e.g., JPEG compr…

physics.geo-ph2022

Deep Neural Networks for Creating Reliable PmP Database with a Case Study in Southern California

Wen Ding, Tianjue Li, Xu Yang +2

Recent progresses in artificial intelligence and machine learning make it possible to automatically identify seismic phases from exponentially growing seismic data. Despite some ex…

cs.CR2020

A Framework for Behavior Privacy Preserving in Radio Frequency Signal

Jianwei Liu, Jinsong Han, Lei Yang +3

Recent years have witnessed the bloom development of the human-centered wireless sensing applications, in which some human information, such as the user's identity and motions, can…

math.NA2023

Error Analysis for the Implicit Boundary Integral Method

Yimin Zhong, Kui Ren, Olof Runborg +1

The implicit boundary integral method (IBIM) provides a framework to construct quadrature rules on regular lattices for integrals over irregular domain boundaries. This work provid…

cs.CV2026

RBD: A Reconstruction-Based Method for Generalizable and Efficient Detection of Fake Images

Qingyu Liu, Zhongjie Ba, Jianmin Guo +4

Recently, reconstruction-based methods have gained attention for AIGC image detection. These methods leverage pre-trained diffusion models to reconstruct inputs and measure residua…

cs.CV2025

Morphology-optimized Multi-Scale Fusion: Combining Local Artifacts and Mesoscopic Semantics for Deepfake Detection and Localization

Chao Shuai, Gaojian Wang, Kun Pan +7

While the pursuit of higher accuracy in deepfake detection remains a central goal, there is an increasing demand for precise localization of manipulated regions. Despite the remark…

cs.SE2023

Demystifying Compiler Unstable Feature Usage and Impacts in the Rust Ecosystem

Chenghao Li, Yifei Wu, Wenbo Shen +5

Rust programming language is gaining popularity rapidly in building reliable and secure systems due to its security guarantees and outstanding performance. To provide extra functio…

cs.CR2021

e-PoS: Making Proof-of-Stake Decentralized and Fair

Muhammad Saad, Zhan Qin, Kui Ren +2

Blockchain applications that rely on the Proof-of-Work (PoW) have increasingly become energy inefficient with a staggering carbon footprint. In contrast, energy-efficient alternati…

cs.CR2024

FedTracker: Furnishing Ownership Verification and Traceability for Federated Learning Model

Shuo Shao, Wenyuan Yang, Hanlin Gu +4

Federated learning (FL) is a distributed machine learning paradigm allowing multiple clients to collaboratively train a global model without sharing their local data. However, FL e…

cs.CR2018

Towards Differentially Private Truth Discovery for Crowd Sensing Systems

Yaliang Li, Houping Xiao, Zhan Qin +5

Nowadays, crowd sensing becomes increasingly more popular due to the ubiquitous usage of mobile devices. However, the quality of such human-generated sensory data varies significan…

cs.CR2025

Textual Unlearning Gives a False Sense of Unlearning

Jiacheng Du, Zhibo Wang, Jie Zhang +3

Language Models (LMs) are prone to ''memorizing'' training data, including substantial sensitive user information. To mitigate privacy risks and safeguard the right to be forgotten…

cs.CR2020

Towards Efficiently Establishing Mutual Distrust Between Host Application and Enclave for SGX

Yuan Chen, Jiaqi Li, Guorui Xu +4

Since its debut, SGX has been used in many applications, e.g., secure data processing. However, previous systems usually assume a trusted enclave and ignore the security issues cau…

eess.SP2020

Injecting Reliable Radio Frequency Fingerprints Using Metasurface for The Internet of Things

Sekhar Rajendran, Zhi Sun, Feng Lin +1

In Internet of Things, where billions of devices with limited resources are communicating with each other, security has become a major stumbling block affecting the progress of thi…

cs.CV2022

Feature Importance-aware Transferable Adversarial Attacks

Zhibo Wang, Hengchang Guo, Zhifei Zhang +3

Transferability of adversarial examples is of central importance for attacking an unknown model, which facilitates adversarial attacks in more practical scenarios, e.g., black-box…

cs.CR2021

Android HIV: A Study of Repackaging Malware for Evading Machine-Learning Detection

Xiao Chen, Chaoran Li, Derui Wang +5

Machine learning based solutions have been successfully employed for automatic detection of malware on Android. However, machine learning models lack robustness to adversarial exam…

cs.CR2026

DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection

Weiwei Qi, Zefeng Wu, Zhilin Guo +5

Most existing LLM safety evaluation and defense methods follow a static formulation: jailbreak vulnerabilities are evaluated with fixed attack methods, and guardrails are trained o…

cs.LG2019

Data Poisoning Attack against Knowledge Graph Embedding

Hengtong Zhang, Tianhang Zheng, Jing Gao +4

Knowledge graph embedding (KGE) is a technique for learning continuous embeddings for entities and relations in the knowledge graph.Due to its benefit to a variety of downstream ta…

cs.CV2026

When Detectors Forget Forensics: Blocking Semantic Shortcuts for Generalizable AI-Generated Image Detection

Chao Shuai, Shaojing Fan, Chenlin Zou +6

The growing realism of generative models has blurred the boundary between real and synthetic content, posing significant challenges to reliable AI-generated image detection. Althou…

cs.CL2026

APPSI-139: A Parallel Corpus of English Application Privacy Policy Summarization and Interpretation

Pengyun Zhu, Qiheng Sun, Long Wen +7

Privacy policies are essential for users to understand how service providers handle their personal data. However, these documents are often long and complex, as well as filled with…

cs.CR2024

ERASER: Machine Unlearning in MLaaS via an Inference Serving-Aware Approach

Yuke Hu, Jian Lou, Jiaqi Liu +4

Over the past years, Machine Learning-as-a-Service (MLaaS) has received a surging demand for supporting Machine Learning-driven services to offer revolutionized user experience acr…

math.CO2014

The Asymptotics of Large Constrained Graphs

Charles Radin, Kui Ren, Lorenzo Sadun

We show, through local estimates and simulation, that if one constrains simple graphs by their densities of edges and of triangles, then asymptotically (in the n…

cs.CR2021

Revisiting Challenges for Selective Data Protection of Real Applications

Lin Ma, Jinyan Xu, Jiadong Sun +5

Selective data protection is a promising technique to defend against the data leakage attack. In this paper, we revisit technical challenges that were neglected when applying this…

cs.CR2024

Releasing Malevolence from Benevolence: The Menace of Benign Data on Machine Unlearning

Binhao Ma, Tianhang Zheng, Hongsheng Hu +5

Machine learning models trained on vast amounts of real or synthetic data often achieve outstanding predictive performance across various domains. However, this utility comes with…

math.CO2017

Surface effects in dense random graphs with sharp edge constraint

Charles Radin, Kui Ren, Lorenzo Sadun

We show that the random number of triangles in a random graph on vertices, with a strict constraint on the total number of edges, admits an expansion $T_n = an^3 + bn^2 +…

cs.CV2025

FSFM: A Generalizable Face Security Foundation Model via Self-Supervised Facial Representation Learning

Gaojian Wang, Feng Lin, Tong Wu +3

This work asks: with abundant, unlabeled real faces, how to learn a robust and transferable facial representation that boosts various face security tasks with respect to generaliza…

physics.optics2025

Phase retrieval via media diversity

Yan Cheng, Kui Ren, Nathan Soedjak

This work studies phase retrieval for wave fields, aiming to recover the phase of an incoming wave from multi-plane intensity measurements behind different types of linear and nonl…

cs.CR2025

Explainer-guided Targeted Adversarial Attacks against Binary Code Similarity Detection Models

Mingjie Chen, Tiancheng Zhu, Mingxue Zhang +4

Binary code similarity detection (BCSD) serves as a fundamental technique for various software engineering tasks, e.g., vulnerability detection and classification. Attacks against…

cs.CR2025

Can Small Language Models Reliably Resist Jailbreak Attacks? A Comprehensive Evaluation

Wenhui Zhang, Huiyu Xu, Zhibo Wang +3

Small language models (SLMs) have emerged as promising alternatives to large language models (LLMs) due to their low computational demands, enhanced privacy guarantees, and compara…

cs.CR2025

Taught Well Learned Ill: Towards Distillation-conditional Backdoor Attack

Yukun Chen, Boheng Li, Yu Yuan +5

Knowledge distillation (KD) is a vital technique for deploying deep neural networks (DNNs) on resource-constrained devices by transferring knowledge from large teacher models to li…

cs.CV2024

Do As I Do: Pose Guided Human Motion Copy

Sifan Wu, Zhenguang Liu, Beibei Zhang +4

Human motion copy is an intriguing yet challenging task in artificial intelligence and computer vision, which strives to generate a fake video of a target person performing the mot…

cs.LG2024

Sora Detector: A Unified Hallucination Detection for Large Text-to-Video Models

Zhixuan Chu, Lei Zhang, Yichen Sun +4

The rapid advancement in text-to-video (T2V) generative models has enabled the synthesis of high-fidelity video content guided by textual descriptions. Despite this significant pro…

cs.CR2026

"Training robust watermarking model may hurt authentication!'' Exploring and Mitigating the Identity Leakage in Robust Watermarking

Xinyu Zhang, Ziping Dong, Qingyu Liu +3

The rapid advancement of generative AI has underscored the critical need for identifying image ownership and protecting copyrights. This makes post-processing image watermarking an…

cs.LG2024

Certified Minimax Unlearning with Generalization Rates and Deletion Capacity

Jiaqi Liu, Jian Lou, Zhan Qin +1

We study the problem of -certified machine unlearning for minimax models. Most of the existing works focus on unlearning from standard statistical learning models that hav…

cs.CR2026

STEP: Detecting Audio Backdoor Attacks via Stability-based Trigger Exposure Profiling

Kun Wang, Meng Chen, Junhao Wang +6

With the widespread deployment of deep-learning-based speech models in security-critical applications, backdoor attacks have emerged as a serious threat: an adversary who poisons a…

cs.CL2025

ContextCache: Context-Aware Semantic Cache for Multi-Turn Queries in Large Language Models

Jianxin Yan, Wangze Ni, Lei Chen +4

Semantic caching significantly reduces computational costs and improves efficiency by storing and reusing large language model (LLM) responses. However, existing systems rely prima…

cs.CR2026

Towards Identification and Intervention of Safety-Critical Parameters in Large Language Models

Weiwei Qi, Zefeng Wu, Tianhang Zheng +4

Ensuring Large Language Model (LLM) safety is crucial, yet the lack of a clear understanding about safety mechanisms hinders the development of precise and reliable methodologies f…

cs.LG2025

Towards Real-world Debiasing: Rethinking Evaluation, Challenge, and Solution

Peng Kuang, Zhibo Wang, Zhixuan Chu +2

Spurious correlations in training data significantly hinder the generalization capability of machine learning models when faced with distribution shifts, leading to the proposition…

cs.LG2025

Perplexity-Aware Data Scaling Law: Perplexity Landscapes Predict Performance for Continual Pre-training

Lei Liu, Hao Zhu, Yue Shen +4

Continual Pre-training (CPT) serves as a fundamental approach for adapting foundation models to domain-specific applications. Scaling laws for pre-training define a power-law relat…

cs.CR2025

MOVE: Effective and Harmless Ownership Verification via Embedded External Features

Yiming Li, Linghui Zhu, Xiaojun Jia +5

Currently, deep neural networks (DNNs) are widely adopted in different applications. Despite its commercial values, training a well-performing DNN is resource-consuming. Accordingl…

cs.CR2025

DualBreach: Efficient Dual-Jailbreaking via Target-Driven Initialization and Multi-Target Optimization

Xinzhe Huang, Kedong Xiu, Tianhang Zheng +5

Recent research has focused on exploring the vulnerabilities of Large Language Models (LLMs), aiming to elicit harmful and/or sensitive content from LLMs. However, due to the insuf…

cs.CV2019

PointCloud Saliency Maps

Tianhang Zheng, Changyou Chen, Junsong Yuan +2

3D point-cloud recognition with PointNet and its variants has received remarkable progress. A missing ingredient, however, is the ability to automatically evaluate point-wise impor…

cs.CV2023

Locate and Verify: A Two-Stream Network for Improved Deepfake Detection

Chao Shuai, Jieming Zhong, Shuang Wu +6

Deepfake has taken the world by storm, triggering a trust crisis. Current deepfake detection methods are typically inadequate in generalizability, with a tendency to overfit to ima…

cs.SE2023

Enabling Runtime Verification of Causal Discovery Algorithms with Automated Conditional Independence Reasoning (Extended Version)

Pingchuan Ma, Zhenlan Ji, Peisen Yao +2

Causal discovery is a powerful technique for identifying causal relationships among variables in data. It has been widely used in various applications in software engineering. Caus…

cs.LG2025

CoKV: Optimizing KV Cache Allocation via Cooperative Game

Qiheng Sun, Hongwei Zhang, Haocheng Xia +3

Large language models (LLMs) have achieved remarkable success on various aspects of human life. However, one of the major challenges in deploying these models is the substantial me…

math.AP2015

Inverse transport problems in quantitative PAT for molecular imaging

Kui Ren, Rongting Zhang, Yimin Zhong

Fluorescence photoacoustic tomography (fPAT) is a molecular imaging modality that combines photoacoustic tomography (PAT) with fluorescence imaging to obtain high-resolution imagin…

cs.AI2026

QCFuse: Query-Aware Cache Fusion via Compressed View for Efficient RAG Serving

Jianxin Yan, Wangze Ni, Zhenxin Li +8

Retrieval-augmented generation (RAG) improves large language model (LLM) answer quality by grounding generation in external evidence, but processing retrieved contexts makes the pr…

cs.CR2026

LoRA-Key: User-Centric LoRA Watermarking for Text-to-Image Diffusion Models

Yaopeng Wang, Qingliang Wang, Zhibo Wang +5

Low-Rank Adaptation (LoRA) has become a widely used mechanism for customizing text-to-image diffusion models, enabling lightweight modules that are shared, reused, and commercializ…

cs.CR2021

Automatically Locating ARM Instructions Deviation between Real Devices and CPU Emulators

Muhui Jiang, Tianyi Xu, Yajin Zhou +5

Emulator is widely used to build dynamic analysis frameworks due to its fine-grained tracing capability, full system monitoring functionality, and scalability of running on differe…

cs.CR2021

DeFiRanger: Detecting Price Manipulation Attacks on DeFi Applications

Siwei Wu, Dabao Wang, Jianting He +5

The rapid growth of Decentralized Finance (DeFi) boosts the Ethereum ecosystem. At the same time, attacks towards DeFi applications (apps) are increasing. However, to the best of o…

cs.CV2026

Ghosts Beneath Textures: Texture-Relation Cues for Cross-Paradigm AI-Generated Image Detection

Haoyu Wang, Yiming Qin, Zhongjie Ba +4

AI-generated images have proliferated rapidly, motivating extensive research. Most existing AI-generated image detectors are developed and evaluated under image-free generation par…

cs.NI2024

Cora: Accelerating Stateful Network Applications with SmartNICs

Shaoke Xi, Jiaqi Gao, Mengqi Liu +7

With the growing performance requirements on networked applications, there is a new trend of offloading stateful network applications to SmartNICs to improve performance and reduce…

cs.LG2026

Instance-Wise Adaptive Sampling for Dataset Construction in Approximating Inverse Problem Solutions

Jiequn Han, Kui Ren, Nathan Soedjak

We propose an instance-wise adaptive sampling framework for constructing compact and informative training datasets for supervised learning of inverse problem solutions. Typical lea…

cs.CR2023

Private Data Valuation and Fair Payment in Data Marketplaces

Zhihua Tian, Jian Liu, Jingyu Li +5

Data valuation is an essential task in a data marketplace. It aims at fairly compensating data owners for their contribution. There is increasing recognition in the machine learnin…

cs.LG2022

OpBoost: A Vertical Federated Tree Boosting Framework Based on Order-Preserving Desensitization

Xiaochen Li, Yuke Hu, Weiran Liu +5

Vertical Federated Learning (FL) is a new paradigm that enables users with non-overlapping attributes of the same data samples to jointly train a model without directly sharing the…

cs.CR2024

A Certified Robust Watermark For Large Language Models

Xianheng Feng, Jian Liu, Kui Ren +1

The effectiveness of watermark algorithms in AI-generated text identification has garnered significant attention. Concurrently, an increasing number of watermark algorithms have be…

cs.CR2024

TabularMark: Watermarking Tabular Datasets for Machine Learning

Yihao Zheng, Haocheng Xia, Junyuan Pang +5

Watermarking is broadly utilized to protect ownership of shared data while preserving data utility. However, existing watermarking methods for tabular datasets fall short on the de…

cs.IT2016

Capacity-achieving and Flicker-free FEC coding scheme for Dimmable Visible Light Communication Based on Polar Codes

Junbin Fang, Zhen Che, Xiaolong Yu +5

Visible light communication (VLC) could provide short-range optical wireless communication together with illumination using LED lighting. However, conventional forward error correc…

cs.LG2022

Purifier: Defending Data Inference Attacks via Transforming Confidence Scores

Ziqi Yang, Lijin Wang, Da Yang +5

Neural networks are susceptible to data inference attacks such as the membership inference attack, the adversarial model inversion attack and the attribute inference attack, where…

cs.AI2025

Towards Evaluation for Real-World LLM Unlearning

Ke Miao, Yuke Hu, Xiaochen Li +4

This paper analyzes the limitations of existing unlearning evaluation metrics in terms of practicality, exactness, and robustness in real-world LLM unlearning scenarios. To overcom…

cs.LG2024

Sampling with Adaptive Variance for Multimodal Distributions

Björn Engquist, Kui Ren, Yunan Yang

We propose and analyze a class of adaptive sampling algorithms for multimodal distributions on a bounded domain, which share a structural resemblance to the classic overdamped Lang…

cs.CL2024

A Causal Explainable Guardrails for Large Language Models

Zhixuan Chu, Yan Wang, Longfei Li +3

Large Language Models (LLMs) have shown impressive performance in natural language tasks, but their outputs can exhibit undesirable attributes or biases. Existing methods for steer…

cs.CR2024

RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent

Huiyu Xu, Wenhui Zhang, Zhibo Wang +5

Recently, advanced Large Language Models (LLMs) such as GPT-4 have been integrated into many real-world applications like Code Copilot. These applications have significantly expand…

cs.CR2020

ARM Pointer Authentication based Forward-Edge and Backward-Edge Control Flow Integrity for Kernels

Yutian Yang, Songbo Zhu, Wenbo Shen +3

Code reuse attacks are still big threats to software and system security. Control flow integrity is a promising technique to defend against such attacks. However, its effectiveness…

cs.LG2026

A Causal Perspective for Enhancing Jailbreak Attack and Defense

Licheng Pan, Yunsheng Lu, Jiexi Liu +5

Uncovering the mechanisms behind "jailbreaks" in large language models (LLMs) is crucial for enhancing their safety and reliability, yet these mechanisms remain poorly understood.…

cs.CR2019

Towards a First Step to Understand the Cryptocurrency Stealing Attack on Ethereum

Zhen Cheng, Xinrui Hou, Runhuai Li +4

We performed the first systematic study of a new attack on Ethereum that steals cryptocurrencies. The attack is due to the unprotected JSON-RPC endpoints existed in Ethereum nodes…

cs.CL2026

HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment

Langqi Yang, Tianhang Zheng, Yixuan Chen +6

The potential of large language models (LLMs) to generate harmful content poses a significant safety risk for data management, as LLMs are increasingly being used as engines for da…

cs.CR2026

TRACE: Task-Aware Adaptive Self-Evolving Agentic Jailbreaking

Churui Zeng, Weiwei Qi, Kedong Xiu +5

The rise of LLM agents introduces a new threat by enabling planning, coding, and even end-to-end execution of expert-level attack workflows. However, this threat remains underexplo…

math.AP2019

Imaging point sources in heterogeneous environments

Kui Ren, Yimin Zhong

Imaging point sources in heterogeneous environments from boundary or far-field measurements has been extensively studied in the past. In most existing results, the environment, rep…

cs.CR2025

FederBoost: Private Federated Learning for GBDT

Zhihua Tian, Rui Zhang, Xiaoyang Hou +4

Federated Learning (FL) has been an emerging trend in machine learning and artificial intelligence. It allows multiple participants to collaboratively train a better global model a…

cs.CV2023

RemovalNet: DNN Fingerprint Removal Attacks

Hongwei Yao, Zheng Li, Kunzhe Huang +3

With the performance of deep neural networks (DNNs) remarkably improving, DNNs have been widely used in many areas. Consequently, the DNN model has become a valuable asset, and its…

cs.CR2025

WMCopier: Forging Invisible Image Watermarks on Arbitrary Images

Ziping Dong, Chao Shuai, Zhongjie Ba +4

Invisible Image Watermarking is crucial for ensuring content provenance and accountability in generative AI. While Gen-AI providers are increasingly integrating invisible watermark…

math.NA2025

A discontinuous Galerkin method for one-dimensional nonlocal wave problems

Qiang Du, Kui Ren, Lu Zhang +1

This paper presents a fully discrete numerical scheme for one-dimensional nonlocal wave equations and provides a rigorous theoretical analysis. To facilitate the spatial discretiza…

cs.CR2025

Towards Label-Only Membership Inference Attack against Pre-trained Large Language Models

Yu He, Boheng Li, Liu Liu +6

Membership Inference Attacks (MIAs) aim to predict whether a data sample belongs to the model's training set or not. Although prior research has extensively explored MIAs in Large…

cs.LG2026

DaDaDa: A Dataset for Data Pricing in Data Marketplaces

Qiheng Sun, Hongwei Zhang, Junxu Liu +4

High-quality data drives machine learning advances across industries. Recognizing the value of data, data transactions are increasingly common, giving rise to many data marketplace…

cs.CV2025

Robust Representation Consistency Model via Contrastive Denoising

Jiachen Lei, Julius Berner, Jiongxiao Wang +5

Robustness is essential for deep neural networks, especially in security-sensitive applications. To this end, randomized smoothing provides theoretical guarantees for certifying ro…

cs.LG2022

Task-aware Similarity Learning for Event-triggered Time Series

Shaoyu Dou, Kai Yang, Yang Jiao +2

Time series analysis has achieved great success in diverse applications such as network security, environmental monitoring, and medical informatics. Learning similarities among dif…

cs.CR2026

LoopTrap: Termination Poisoning Attacks on LLM Agents

Huiyu Xu, Zhibo Wang, Wenhui Zhang +4

Modern LLM agents solve complex tasks by operating in iterative execution loops, where they repeatedly reason, act, and self-evaluate progress to determine when a task is complete.…

cs.PL2026

Synthesizing Best Abstract Transformers via Parallel Bit-Vector Optimization

Weiqi Wang, Peisen Yao, Hanrui Zuo +3

Abstract interpretation provides a principled foundation for constructing sound static analyses through systematic abstraction. A central challenge is synthesizing the best abstrac…

cs.CR2022

iLibScope: Reliable Third-Party Library Detection for iOS Mobile Apps

Jingyi Guo, Min Zheng, Yajin Zhou +4

Vetting security impacts introduced by third-party libraries in iOS apps requires a reliable library detection technique. Especially when a new vulnerability (or a privacy-invasive…

math.AP2017

A global stability estimate for the photo-acoustic inverse problem in layered media

Kui Ren, Faouzi Triki

This paper is concerned with the stability issue in determining absorption and diffusion coefficients in photoacoustic imaging. Assuming that the medium is layered and the acoustic…

cs.CV2024

SurrogatePrompt: Bypassing the Safety Filter of Text-to-Image Models via Substitution

Zhongjie Ba, Jieming Zhong, Jiachen Lei +5

Advanced text-to-image models such as DALLE 2 and Midjourney possess the capacity to generate highly realistic images, raising significant concerns regarding the potential p…

cs.CV2023

ANetQA: A Large-scale Benchmark for Fine-grained Compositional Reasoning over Untrimmed Videos

Zhou Yu, Lixiang Zheng, Zhou Zhao +4

Building benchmarks to systemically analyze different capabilities of video question answering (VideoQA) models is challenging yet crucial. Existing benchmarks often use non-compos…

cs.CV2022

Vanilla Feature Distillation for Improving the Accuracy-Robustness Trade-Off in Adversarial Training

Guodong Cao, Zhibo Wang, Xiaowei Dong +4

Adversarial training has been widely explored for mitigating attacks against deep models. However, most existing works are still trapped in the dilemma between higher accuracy and…

cs.PL2026

A Fresh Look at Best Inductive Loop Invariant Synthesis for Bit-Vector Relations

Hanrui Zuo, Peisen Yao, Kui Ren

The paper proposes a new optimization-based formulation for synthesizing best inductive invariants in bit‑vector programs and introduces two algorithms—a guided linear search and a…

#inductive invariant synthesis#bit-vector verification#formal methods#optimization
cs.CL2025

Mitigating Social Bias in Large Language Models: A Multi-Objective Approach within a Multi-Agent Framework

Zhenjie Xu, Wenqing Chen, Yi Tang +6

Natural language processing (NLP) has seen remarkable advancements with the development of large language models (LLMs). Despite these advancements, LLMs often produce socially bia…

cs.CV2023

Action Recognition with Multi-stream Motion Modeling and Mutual Information Maximization

Yuheng Yang, Haipeng Chen, Zhenguang Liu +5

Action recognition has long been a fundamental and intriguing problem in artificial intelligence. The task is challenging due to the high dimensionality nature of an action, as wel…

cs.CL2026

Explainable Token-level Noise Filtering for LLM Fine-tuning Datasets

Yuchen Yang, Wenze Lin, Enhao Huang +6

Large Language Models (LLMs) have seen remarkable advancements, achieving state-of-the-art results in diverse applications. Fine-tuning, an important step for adapting LLMs to spec…

cs.CR2024

SoK: On the Role and Future of AIGC Watermarking in the Era of Gen-AI

Kui Ren, Ziqi Yang, Li Lu +6

The rapid advancement of AI technology, particularly in generating AI-generated content (AIGC), has transformed numerous fields, e.g., art video generation, but also brings new ris…

cs.PL2026

Accelerating C/C++ Pointer Analysis via Compiler-Based Offline Simplifications

Zinan Gu, Peisen Yao, Kui Ren

Pointer analysis is a cornerstone of numerous static analysis applications, including compiler optimizations, slicing, bug detection, and verification. While offline simplification…

cs.CR2022

Backdoor Defense via Decoupling the Training Process

Kunzhe Huang, Yiming Li, Baoyuan Wu +2

Recent studies have revealed that deep neural networks (DNNs) are vulnerable to backdoor attacks, where attackers embed hidden backdoors in the DNN model by poisoning a few trainin…

math.NA2016

Numerical algorithms based on Galerkin methods for the modeling of reactive interfaces in photoelectrochemical (PEC) solar cells

Michael Harmon, Irene M. Gamba, Kui Ren

This work concerns the numerical solution of a coupled system of self-consistent reaction-drift-diffusion-Poisson equations that describes the macroscopic dynamics of charge transp…

cs.CR2021

A Measurement Study on the (In)security of End-of-Life (EoL) Embedded Devices

Dingding Wang, Muhui Jiang, Rui Chang +5

Embedded devices are becoming popular. Meanwhile, researchers are actively working on improving the security of embedded devices. However, previous work ignores the insecurity caus…

math.NA2020

The quadratic Wasserstein metric for inverse data matching

Bjorn Engquist, Kui Ren, Yunan Yang

This work characterizes, analytically and numerically, two major effects of the quadratic Wasserstein () distance as the measure of data discrepancy in computational solutions…

cs.CV2025

Harnessing Frequency Spectrum Insights for Image Copyright Protection Against Diffusion Models

Zhenguang Liu, Chao Shuai, Shaojing Fan +4

Diffusion models have achieved remarkable success in novel view synthesis, but their reliance on large, diverse, and often untraceable Web datasets has raised pressing concerns abo…