papers

Publications (69)

cs.CL2025

Turbocharging Web Automation: The Impact of Compressed History States

Xiyue Zhu, Peng Tang, Haofu Liao +1

Language models have led to a leap forward in web automation. The current web automation approaches take the current web state, history actions, and language instruction as inputs…

cs.LG2024

Enhanced Traffic Flow Prediction with Multi-Segment Fusion Tensor Graph Convolutional Networks

Wei Zhang, Peng Tang

Accurate traffic Flow Prediction can assist in traffic management, route planning, and congestion mitigation, which holds significant importance in enhancing the efficiency and rel…

cs.CV2020

Look Closer to Ground Better: Weakly-Supervised Temporal Grounding of Sentence in Video

Zhenfang Chen, Lin Ma, Wenhan Luo +2

In this paper, we study the problem of weakly-supervised temporal grounding of sentence in video. Specifically, given an untrimmed video and a query sentence, our goal is to locali…

cs.LG2026

Learning to Explore: Policy-Guided Outlier Synthesis for Graph Out-of-Distribution Detection

Li Sun, Lanxu Yang, Jiayu Tian +6

Detecting out-of-distribution (OOD) graphs is crucial for ensuring the safety and reliability of Graph Neural Networks. In unsupervised graph-level OOD detection, models are typica…

cs.CL2026

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments

Deyao Zhu, Xin Zhou, Shengling Qin +44

Pretraining scaling laws reveal that model capability improves predictably with data and compute. But learning from real world environments after deployment remains far less unders…

cs.LG2023

RNN-Guard: Certified Robustness Against Multi-frame Attacks for Recurrent Neural Networks

Yunruo Zhang, Tianyu Du, Shouling Ji +2

It is well-known that recurrent neural networks (RNNs), although widely used, are vulnerable to adversarial attacks including one-frame attacks and multi-frame attacks. Though a fe…

eess.IV2022

Unsupervised Tissue Segmentation via Deep Constrained Gaussian Network

Yang Nan, Peng Tang, Guyue Zhang +5

Tissue segmentation is the mainstay of pathological examination, whereas the manual delineation is unduly burdensome. To assist this time-consuming and subjective manual step, rese…

astro-ph.IM2018

Design of Near Infrared Sky Brightness Monitor and Test Running at Ngari Observatory in Tibet

Qi-Jie Tang, Jian Wang, Shu-cheng Dong +16

Tibet is known as the third pole of the earth, as high as the South Pole and North Pole. The Ngari (Ali) observatory in Tibet has the advantage of plenty of photometric night, low…

cs.CV2026

CLEAR: Context-Aware Learning with End-to-End Mask-Free Inference for Adaptive Video Subtitle Removal

Qingdong He, Chaoyi Wang, Peng Tang +2

Video subtitle removal aims to distinguish text overlays from background content while preserving temporal coherence. Existing diffusion-based methods necessitate explicit mask seq…

cs.AI2026

ShortcutBreaker: Low-Rank Noisy Bottleneck and Frequency Filtering Block for Multi-Class Unsupervised Anomaly Detection

Peng Tang, Xiaobin Hu, Tingcheng Li +3

Multi-class unsupervised anomaly detection (MUAD) has garnered growing research interest, as it seeks to develop a unified model for anomaly detection across multiple classes, i.e.…

cs.LG2025

MoE-SpeQ: Speculative Quantized Decoding with Proactive Expert Prefetching and Offloading for Mixture-of-Experts

Wenfeng Wang, Jiacheng Liu, Xiaofeng Hou +5

The immense memory requirements of state-of-the-art Mixture-of-Experts (MoE) models present a significant challenge for inference, often exceeding the capacity of a single accelera…

cs.IR2024

SSDRec: Self-Augmented Sequence Denoising for Sequential Recommendation

Chi Zhang, Qilong Han, Rui Chen +3

Traditional sequential recommendation methods assume that users' sequence data is clean enough to learn accurate sequence representations to reflect user preferences. In practice,…

eess.IV2022

Fuzzy Attention Neural Network to Tackle Discontinuity in Airway Segmentation

Yang Nan, Javier Del Ser, Zeyu Tang +7

Airway segmentation is crucial for the examination, diagnosis, and prognosis of lung diseases, while its manual delineation is unduly burdensome. To alleviate this time-consuming a…

cs.DC2019

Security modeling and efficient computation offloading for service workflow in mobile edge computing

Binbin Huang, Zhongjin Lia, Peng Tang +5

It is a big challenge for resource-limited mobile devices (MDs) to execute various complex and energy-consumed mobile applications. Fortunately, as a novel computing paradigm, edge…

cs.CV2025

R-VLM: Region-Aware Vision Language Model for Precise GUI Grounding

Joonhyung Park, Peng Tang, Sagnik Das +4

Visual agent models for automating human activities on Graphical User Interfaces (GUIs) have emerged as a promising research direction, driven by advances in large Vision Language…

cs.CV2021

Learning Inductive Attention Guidance for Partially Supervised Pancreatic Ductal Adenocarcinoma Prediction

Yan Wang, Peng Tang, Yuyin Zhou +3

Pancreatic ductal adenocarcinoma (PDAC) is the third most common cause of cancer death in the United States. Predicting tumors like PDACs (including both classification and segment…

cs.CV2017

Deep Patch Learning for Weakly Supervised Object Classification and Discovery

Peng Tang, Xinggang Wang, Zilong Huang +2

Patch-level image representation is very important for object classification and detection, since it is robust to spatial transformation, scale variation, and cluttered background.…

cs.CV2023

Graph-Ensemble Learning Model for Multi-label Skin Lesion Classification using Dermoscopy and Clinical Images

Peng Tang, Yang Nan, Tobias Lasser

Many skin lesion analysis (SLA) methods recently focused on developing a multi-modal-based multi-label classification method due to two factors. The first is multi-modal data, i.e.…

cs.CV2024

Synthesize Step-by-Step: Tools, Templates and LLMs as Data Generators for Reasoning-Based Chart VQA

Zhuowan Li, Bhavan Jasani, Peng Tang +1

Understanding data visualizations like charts and plots requires reasoning about both visual elements and numerics. Although strong in extractive questions, current chart visual qu…

cs.CV2025

TokenAR: Multiple Subject Generation via Autoregressive Token-level enhancement

Haiyue Sun, Qingdong He, Jinlong Peng +5

Autoregressive Model (AR) has shown remarkable success in conditional image generation. However, these approaches for multiple reference generation struggle with decoupling differe…

cs.CV2018

Training Multi-organ Segmentation Networks with Sample Selection by Relaxed Upper Confident Bound

Yan Wang, Yuyin Zhou, Peng Tang +3

Deep convolutional neural networks (CNNs), especially fully convolutional networks, have been widely applied to automatic medical image segmentation problems, e.g., multi-organ seg…

cs.CV2023

SR-RKAC: Improving Single Image Defocus Deblurring

Peng Tang, Zhiqiang Xu, Pengfei Wei +5

We propose an efficient deep learning method for single image defocus deblurring (SIDD) by further exploring inverse kernel properties. Although the current inverse kernel method,…

cs.LG2025

Devil's Hand: Data Poisoning Attacks to Locally Private Graph Learning Protocols

Longzhu He, Chaozhuo Li, Peng Tang +3

Graph neural networks (GNNs) have achieved significant success in graph representation learning and have been applied to various domains. However, many real-world graphs contain se…

eess.IV2024

Federated Semi-supervised Learning for Medical Image Segmentation with intra-client and inter-client Consistency

Yubin Zheng, Peng Tang, Tianjie Ju +2

Medical image segmentation plays a vital role in clinic disease diagnosis and medical image analysis. However, labeling medical images for segmentation task is tough due to the ind…

cs.PF2023

MMBench: Benchmarking End-to-End Multi-modal DNNs and Understanding Their Hardware-Software Implications

Cheng Xu, Xiaofeng Hou, Jiacheng Liu +9

The explosive growth of various types of big data and advances in AI technologies have catalyzed a new type of workloads called multi-modal DNNs. Multi-modal DNNs are capable of in…

cs.CV2023

Multiple-Question Multiple-Answer Text-VQA

Peng Tang, Srikar Appalaraju, R. Manmatha +2

We present Multiple-Question Multiple-Answer (MQMA), a novel approach to do text-VQA in encoder-decoder transformer models. The text-VQA task requires a model to answer a question…

cs.CR2024

PassTSL: Modeling Human-Created Passwords through Two-Stage Learning

Yangde Wang, Haozhang Li, Weidong Qiu +2

Textual passwords are still the most widely used user authentication mechanism. Due to the close connections between textual passwords and natural languages, advanced technologies…

cs.CL2025

Using LLMs for Automated Privacy Policy Analysis: Prompt Engineering, Fine-Tuning and Explainability

Yuxin Chen, Peng Tang, Weidong Qiu +1

Privacy policies are widely used by digital services and often required for legal purposes. Many machine learning based classifiers have been developed to automate detection of dif…

cs.LG2025

A Survey on Inference Optimization Techniques for Mixture of Experts Models

Jiacheng Liu, Peng Tang, Wenfeng Wang +5

The emergence of large-scale Mixture of Experts (MoE) models represents a significant advancement in artificial intelligence, offering enhanced model capacity and computational eff…

cs.CV2023

Joint-Individual Fusion Structure with Fusion Attention Module for Multi-Modal Skin Cancer Classification

Peng Tang, Xintong Yan, Yang Nan +3

Most convolutional neural network (CNN) based methods for skin cancer classification obtain their results using only dermatological images. Although good classification results hav…

cs.LG2024

HOBBIT: A Mixed Precision Expert Offloading System for Fast MoE Inference

Peng Tang, Jiacheng Liu, Xiaofeng Hou +5

The Mixture-of-Experts (MoE) architecture has demonstrated significant advantages in the era of Large Language Models (LLMs), offering enhanced capabilities with reduced inference…

cs.CV2018

Rethink ReLU to Training Better CNNs

Gangming Zhao, Zhaoxiang Zhang, He Guan +2

Most of convolutional neural networks share the same characteristic: each convolutional layer is followed by a nonlinear activation layer where Rectified Linear Unit (ReLU) is the…

cs.CR2025

Poisoning Attacks to Local Differential Privacy for Ranking Estimation

Pei Zhan, Peng Tang, Yangzhuo Li +2

Local differential privacy (LDP) involves users perturbing their inputs to provide plausible deniability of their data. However, this also makes LDP vulnerable to poisoning attacks…

cs.LG2025

FedTopo: Topology-Informed Representation Alignment in Federated Learning under Non-I.I.D. Conditions

Ke Hu, Liyao Xiang, Peng Tang +1

Current federated-learning models deteriorate under heterogeneous (non-I.I.D.) client data, as their feature representations diverge and pixel- or patch-level objectives fail to ca…

cs.CV2018

Semi-Supervised Multi-Organ Segmentation via Deep Multi-Planar Co-Training

Yuyin Zhou, Yan Wang, Peng Tang +4

In multi-organ segmentation of abdominal CT scans, most existing fully supervised deep learning algorithms require lots of voxel-wise annotations, which are usually difficult, expe…

cs.CR2026

A Comprehensive Study on GDPR-Oriented Analysis of Privacy Policies: Taxonomy, Corpus and GDPR Concept Classifiers

Peng Tang, Xin Li, Yuxin Chen +5

Machine learning based classifiers that take a privacy policy as the input and predict relevant concepts are useful in different applications such as (semi-)automated compliance an…

cs.SE2025

OpenDerisk: An Industrial Framework for AI-Driven SRE, with Design, Implementation, and Case Studies

Peng Di, Faqiang Chen, Xiao Bai +12

The escalating complexity of modern software imposes an unsustainable operational burden on Site Reliability Engineering (SRE) teams, demanding AI-driven automation that can emulat…

cs.LG2024

An Adaptive Latent Factorization of Tensors Model for Embedding Dynamic Communication Network

Xin Liao, Qicong Hu, Peng Tang

The Dynamic Communication Network (DCN) describes the interactions over time among various communication nodes, and it is widely used in Big-data applications as a data source. As…

cs.CL2025

MoE-Prism: Disentangling Monolithic Experts for Elastic MoE Services via Model-System Co-Designs

Xinfeng Xia, Jiacheng Liu, Xiaofeng Hou +4

Mixture-of-Experts (MoE) models, the state-of-the-art in large-scale AI, achieve high quality by sparsely activating parameters. However, their reliance on routing between a few mo…

cs.CV2018

PCL: Proposal Cluster Learning for Weakly Supervised Object Detection

Peng Tang, Xinggang Wang, Song Bai +4

Weakly Supervised Object Detection (WSOD), using only image-level annotations to train object detectors, is of growing importance in object recognition. In this paper, we propose a…

eess.IV2024

Single-Shared Network with Prior-Inspired Loss for Parameter-Efficient Multi-Modal Imaging Skin Lesion Classification

Peng Tang, Tobias Lasser

In this study, we introduce a multi-modal approach that efficiently integrates multi-scale clinical and dermoscopy features within a single network, thereby substantially reducing…

cs.LG2026

Towards Personalized Differentially Private Learning for Decentralized Local Graphs

Longzhu He, Peng Tang, Chaozhuo Li +5

Graph-structured data is increasingly generated and stored in decentralized environments, such as social platforms, mobile applications, and edge networks, where users maintain con…

cs.CV2026

Towards Generalized Multi-Image Editing for Unified Multimodal Models

Pengcheng Xu, Peng Tang, Donghao Luo +7

Unified Multimodal Models (UMMs) integrate multimodal understanding and generation, yet they are limited to maintaining visual consistency and disambiguating visual cues when refer…

cs.CV2023

DocFormerv2: Local Features for Document Understanding

Srikar Appalaraju, Peng Tang, Qi Dong +3

We propose DocFormerv2, a multi-modal transformer for Visual Document Understanding (VDU). The VDU domain entails understanding documents (beyond mere OCR predictions) e.g., extrac…

cs.CV2026

FFP-300K: Scaling First-Frame Propagation for Generalizable Video Editing

Xijie Huang, Chengming Xu, Donghao Luo +6

First-Frame Propagation (FFP) offers a promising paradigm for controllable video editing, but existing methods are hampered by a reliance on cumbersome run-time guidance. We identi…

cs.CV2020

Proposal Learning for Semi-Supervised Object Detection

Peng Tang, Chetan Ramaiah, Yan Wang +2

In this paper, we focus on semi-supervised object detection to boost performance of proposal-based object detectors (a.k.a. two-stage object detectors) by training on both labeled…

cs.CV2024

DocKD: Knowledge Distillation from LLMs for Open-World Document Understanding Models

Sungnyun Kim, Haofu Liao, Srikar Appalaraju +6

Visual document understanding (VDU) is a challenging task that involves understanding documents across various modalities (text and image) and layouts (forms, tables, etc.). This s…

cs.CV2019

Object Detection in Videos by High Quality Object Linking

Peng Tang, Chunyu Wang, Xinggang Wang +3

Compared with object detection in static images, object detection in videos is more challenging due to degraded image qualities. An effective way to address this problem is to expl…

physics.optics2026

Self-locking non-volatile coding metasurfaces via origami-based mechanical bits

Ding Zhang, Peng Tang, Liqiao Jing +7

Digital coding metasurfaces have revolutionized electromagnetic (EM) manipulation, yet typical tunable approaches based on active components suffer from the "volatility" bottleneck…

stat.ML2016

Revisiting Multiple Instance Neural Networks

Xinggang Wang, Yongluan Yan, Peng Tang +2

Recently neural networks and multiple instance learning are both attractive topics in Artificial Intelligence related research fields. Deep neural networks have achieved great succ…

cs.LG2023

Feature Norm Regularized Federated Learning: Transforming Skewed Distributions into Global Insights

Ke Hu, WeiDong Qiu, Peng Tang

In the field of federated learning, addressing non-independent and identically distributed (non-i.i.d.) data remains a quintessential challenge for improving global model performan…

stat.ME2021

Robust Estimation of Sparse Precision Matrix using Adaptive Weighted Graphical Lasso Approach

Peng Tang, Huijing Jiang, Heeyoung Kim +1

Estimation of a precision matrix (i.e., inverse covariance matrix) is widely used to exploit conditional independence among continuous variables. The influence of abnormal observat…

cs.DC2026

PCR: A Prefetch-Enhanced Cache Reuse System for Low-Latency RAG Serving

Wenfeng Wang, Xiaofeng Hou, Peng Tang +5

Retrieval-Augmented Generation (RAG) systems enhance the performance of large language models (LLMs) by incorporating supplementary retrieved documents, enabling more accurate and…

cs.CV2023

DEED: Dynamic Early Exit on Decoder for Accelerating Encoder-Decoder Transformer Models

Peng Tang, Pengkai Zhu, Tian Li +3

Encoder-decoder transformer models have achieved great success on various vision-language (VL) tasks, but they suffer from high inference latency. Typically, the decoder takes up m…

eess.IV2024

Pay Less On Clinical Images: Asymmetric Multi-Modal Fusion Method For Efficient Multi-Label Skin Lesion Classification

Peng Tang, Tobias Lasser

Existing multi-modal approaches primarily focus on enhancing multi-label skin lesion classification performance through advanced fusion modules, often neglecting the associated ris…

cs.LG2026

LLMZero: Discovering Adaptive Training Strategies for RL Post-Training via LLM Agents

Haoyang Fang, Wei Zhu, Boran Han +11

RL post-training strategies are dataset-dependent and reveal a recurring empirical pattern: capacity parameters accumulate monotonically across stages, while regularization paramet…

cs.LG2024

MDS-GNN: A Mutual Dual-Stream Graph Neural Network on Graphs with Incomplete Features and Structure

Peng Yuan, Peng Tang

Graph Neural Networks (GNNs) have emerged as powerful tools for analyzing and learning representations from graph-structured data. A crucial prerequisite for the outstanding perfor…

cs.CV2017

Multiple Instance Detection Network with Online Instance Classifier Refinement

Peng Tang, Xinggang Wang, Xiang Bai +1

Of late, weakly supervised object detection is with great importance in object recognition. Based on deep learning, weakly supervised detectors have achieved many promising results…

cs.HC2025

Everyone's Privacy Matters! An Analysis of Privacy Leakage from Real-World Facial Images on Twitter and Associated User Behaviors

Yuqi Niu, Weidong Qiu, Peng Tang +5

Online users often post facial images of themselves and other people on online social networks (OSNs) and other Web 2.0 platforms, which can lead to potential privacy leakage of pe…

cs.CV2026

Open-Vocabulary Semantic Segmentation Network Integrating Object-Level Label and Scene-Level Semantic Features for Multimodal Remote Sensing Images

Jinkun Dai, Yuanxin Ye, Peng Tang +4

Semantic segmentation of multi-modal remote sensing imagery plays a pivotal role in land use/land cover (LULC) mapping, environmental monitoring, and precision earth observation. C…

cs.CV2025

FedDEAP: Adaptive Dual-Prompt Tuning for Multi-Domain Federated Learning

Yubin Zheng, Pak-Hei Yeung, Jing Xia +4

Federated learning (FL) enables multiple clients to collaboratively train machine learning models without exposing local data, balancing performance and privacy. However, domain sh…

cs.CV2026

The devil is in the details: Enhancing Video Virtual Try-On via Keyframe-Driven Details Injection

Qingdong He, Xueqin Chen, Yanjie Pan +7

Although diffusion transformer (DiT)-based video virtual try-on (VVT) has made significant progress in synthesizing realistic videos, existing methods still struggle to capture fin…

eess.IV2022

Automatic Fine-grained Glomerular Lesion Recognition in Kidney Pathology

Yang Nan, Fengyi Li, Peng Tang +5

Recognition of glomeruli lesions is the key for diagnosis and treatment planning in kidney pathology; however, the coexisting glomerular structures such as mesangial regions exacer…

cs.CV2019

Robustness of Object Recognition under Extreme Occlusion in Humans and Computational Models

Hongru Zhu, Peng Tang, Jeongho Park +2

Most objects in the visual world are partially occluded, but humans can recognize them without difficulty. However, it remains unknown whether object recognition models like convol…

cs.LG2025

Multi-Head Self-Attending Neural Tucker Factorization

Yikai Hou, Peng Tang

Quality-of-service (QoS) data exhibit dynamic temporal patterns that are crucial for accurately predicting missing values. These patterns arise from the evolving interactions betwe…

cs.CV2021

Shape-Texture Debiased Neural Network Training

Yingwei Li, Qihang Yu, Mingxing Tan +5

Shape and texture are two prominent and complementary cues for recognizing objects. Nonetheless, Convolutional Neural Networks are often biased towards either texture or shape, dep…

cs.CR2025

SE#PCFG: Semantically Enhanced PCFG for Password Analysis and Cracking

Yangde Wang, Weidong Qiu, Peng Tang +2

Much research has been done on user-generated textual passwords. Surprisingly, semantic information in such passwords remain under-investigated, with passwords created by English-…

cs.LG2025

Neural Canonical Polyadic Factorization for Traffic Analysis

Wenyu Luo, Yikai Hou, Peng Tang

Modern intelligent transportation systems rely on accurate spatiotemporal traffic analysis to optimize urban mobility and infrastructure resilience. However, pervasive missing data…

cs.CV2016

Deep FisherNet for Object Classification

Peng Tang, Xinggang Wang, Baoguang Shi +3

Despite the great success of convolutional neural networks (CNN) for the image classification task on datasets like Cifar and ImageNet, CNN's representation power is still somewhat…