Publications (69)
Turbocharging Web Automation: The Impact of Compressed History States
Xiyue Zhu, Peng Tang, Haofu Liao +1
Language models have led to a leap forward in web automation. The current web automation approaches take the current web state, history actions, and language instruction as inputs…
Enhanced Traffic Flow Prediction with Multi-Segment Fusion Tensor Graph Convolutional Networks
Wei Zhang, Peng Tang
Accurate traffic Flow Prediction can assist in traffic management, route planning, and congestion mitigation, which holds significant importance in enhancing the efficiency and rel…
Look Closer to Ground Better: Weakly-Supervised Temporal Grounding of Sentence in Video
Zhenfang Chen, Lin Ma, Wenhan Luo +2
In this paper, we study the problem of weakly-supervised temporal grounding of sentence in video. Specifically, given an untrimmed video and a query sentence, our goal is to locali…
Learning to Explore: Policy-Guided Outlier Synthesis for Graph Out-of-Distribution Detection
Li Sun, Lanxu Yang, Jiayu Tian +6
Detecting out-of-distribution (OOD) graphs is crucial for ensuring the safety and reliability of Graph Neural Networks. In unsupervised graph-level OOD detection, models are typica…
EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments
Deyao Zhu, Xin Zhou, Shengling Qin +44
Pretraining scaling laws reveal that model capability improves predictably with data and compute. But learning from real world environments after deployment remains far less unders…
RNN-Guard: Certified Robustness Against Multi-frame Attacks for Recurrent Neural Networks
Yunruo Zhang, Tianyu Du, Shouling Ji +2
It is well-known that recurrent neural networks (RNNs), although widely used, are vulnerable to adversarial attacks including one-frame attacks and multi-frame attacks. Though a fe…
Unsupervised Tissue Segmentation via Deep Constrained Gaussian Network
Yang Nan, Peng Tang, Guyue Zhang +5
Tissue segmentation is the mainstay of pathological examination, whereas the manual delineation is unduly burdensome. To assist this time-consuming and subjective manual step, rese…
Design of Near Infrared Sky Brightness Monitor and Test Running at Ngari Observatory in Tibet
Qi-Jie Tang, Jian Wang, Shu-cheng Dong +16
Tibet is known as the third pole of the earth, as high as the South Pole and North Pole. The Ngari (Ali) observatory in Tibet has the advantage of plenty of photometric night, low…
CLEAR: Context-Aware Learning with End-to-End Mask-Free Inference for Adaptive Video Subtitle Removal
Qingdong He, Chaoyi Wang, Peng Tang +2
Video subtitle removal aims to distinguish text overlays from background content while preserving temporal coherence. Existing diffusion-based methods necessitate explicit mask seq…
ShortcutBreaker: Low-Rank Noisy Bottleneck and Frequency Filtering Block for Multi-Class Unsupervised Anomaly Detection
Peng Tang, Xiaobin Hu, Tingcheng Li +3
Multi-class unsupervised anomaly detection (MUAD) has garnered growing research interest, as it seeks to develop a unified model for anomaly detection across multiple classes, i.e.…
MoE-SpeQ: Speculative Quantized Decoding with Proactive Expert Prefetching and Offloading for Mixture-of-Experts
Wenfeng Wang, Jiacheng Liu, Xiaofeng Hou +5
The immense memory requirements of state-of-the-art Mixture-of-Experts (MoE) models present a significant challenge for inference, often exceeding the capacity of a single accelera…
SSDRec: Self-Augmented Sequence Denoising for Sequential Recommendation
Chi Zhang, Qilong Han, Rui Chen +3
Traditional sequential recommendation methods assume that users' sequence data is clean enough to learn accurate sequence representations to reflect user preferences. In practice,…
Fuzzy Attention Neural Network to Tackle Discontinuity in Airway Segmentation
Yang Nan, Javier Del Ser, Zeyu Tang +7
Airway segmentation is crucial for the examination, diagnosis, and prognosis of lung diseases, while its manual delineation is unduly burdensome. To alleviate this time-consuming a…
Security modeling and efficient computation offloading for service workflow in mobile edge computing
Binbin Huang, Zhongjin Lia, Peng Tang +5
It is a big challenge for resource-limited mobile devices (MDs) to execute various complex and energy-consumed mobile applications. Fortunately, as a novel computing paradigm, edge…
R-VLM: Region-Aware Vision Language Model for Precise GUI Grounding
Joonhyung Park, Peng Tang, Sagnik Das +4
Visual agent models for automating human activities on Graphical User Interfaces (GUIs) have emerged as a promising research direction, driven by advances in large Vision Language…
Learning Inductive Attention Guidance for Partially Supervised Pancreatic Ductal Adenocarcinoma Prediction
Yan Wang, Peng Tang, Yuyin Zhou +3
Pancreatic ductal adenocarcinoma (PDAC) is the third most common cause of cancer death in the United States. Predicting tumors like PDACs (including both classification and segment…
Deep Patch Learning for Weakly Supervised Object Classification and Discovery
Peng Tang, Xinggang Wang, Zilong Huang +2
Patch-level image representation is very important for object classification and detection, since it is robust to spatial transformation, scale variation, and cluttered background.…
Graph-Ensemble Learning Model for Multi-label Skin Lesion Classification using Dermoscopy and Clinical Images
Peng Tang, Yang Nan, Tobias Lasser
Many skin lesion analysis (SLA) methods recently focused on developing a multi-modal-based multi-label classification method due to two factors. The first is multi-modal data, i.e.…
Synthesize Step-by-Step: Tools, Templates and LLMs as Data Generators for Reasoning-Based Chart VQA
Zhuowan Li, Bhavan Jasani, Peng Tang +1
Understanding data visualizations like charts and plots requires reasoning about both visual elements and numerics. Although strong in extractive questions, current chart visual qu…
TokenAR: Multiple Subject Generation via Autoregressive Token-level enhancement
Haiyue Sun, Qingdong He, Jinlong Peng +5
Autoregressive Model (AR) has shown remarkable success in conditional image generation. However, these approaches for multiple reference generation struggle with decoupling differe…
Training Multi-organ Segmentation Networks with Sample Selection by Relaxed Upper Confident Bound
Yan Wang, Yuyin Zhou, Peng Tang +3
Deep convolutional neural networks (CNNs), especially fully convolutional networks, have been widely applied to automatic medical image segmentation problems, e.g., multi-organ seg…
SR-RKAC: Improving Single Image Defocus Deblurring
Peng Tang, Zhiqiang Xu, Pengfei Wei +5
We propose an efficient deep learning method for single image defocus deblurring (SIDD) by further exploring inverse kernel properties. Although the current inverse kernel method,…
Devil's Hand: Data Poisoning Attacks to Locally Private Graph Learning Protocols
Longzhu He, Chaozhuo Li, Peng Tang +3
Graph neural networks (GNNs) have achieved significant success in graph representation learning and have been applied to various domains. However, many real-world graphs contain se…
Federated Semi-supervised Learning for Medical Image Segmentation with intra-client and inter-client Consistency
Yubin Zheng, Peng Tang, Tianjie Ju +2
Medical image segmentation plays a vital role in clinic disease diagnosis and medical image analysis. However, labeling medical images for segmentation task is tough due to the ind…
MMBench: Benchmarking End-to-End Multi-modal DNNs and Understanding Their Hardware-Software Implications
Cheng Xu, Xiaofeng Hou, Jiacheng Liu +9
The explosive growth of various types of big data and advances in AI technologies have catalyzed a new type of workloads called multi-modal DNNs. Multi-modal DNNs are capable of in…
Multiple-Question Multiple-Answer Text-VQA
Peng Tang, Srikar Appalaraju, R. Manmatha +2
We present Multiple-Question Multiple-Answer (MQMA), a novel approach to do text-VQA in encoder-decoder transformer models. The text-VQA task requires a model to answer a question…
PassTSL: Modeling Human-Created Passwords through Two-Stage Learning
Yangde Wang, Haozhang Li, Weidong Qiu +2
Textual passwords are still the most widely used user authentication mechanism. Due to the close connections between textual passwords and natural languages, advanced technologies…
Using LLMs for Automated Privacy Policy Analysis: Prompt Engineering, Fine-Tuning and Explainability
Yuxin Chen, Peng Tang, Weidong Qiu +1
Privacy policies are widely used by digital services and often required for legal purposes. Many machine learning based classifiers have been developed to automate detection of dif…
A Survey on Inference Optimization Techniques for Mixture of Experts Models
Jiacheng Liu, Peng Tang, Wenfeng Wang +5
The emergence of large-scale Mixture of Experts (MoE) models represents a significant advancement in artificial intelligence, offering enhanced model capacity and computational eff…
Joint-Individual Fusion Structure with Fusion Attention Module for Multi-Modal Skin Cancer Classification
Peng Tang, Xintong Yan, Yang Nan +3
Most convolutional neural network (CNN) based methods for skin cancer classification obtain their results using only dermatological images. Although good classification results hav…
HOBBIT: A Mixed Precision Expert Offloading System for Fast MoE Inference
Peng Tang, Jiacheng Liu, Xiaofeng Hou +5
The Mixture-of-Experts (MoE) architecture has demonstrated significant advantages in the era of Large Language Models (LLMs), offering enhanced capabilities with reduced inference…
Rethink ReLU to Training Better CNNs
Gangming Zhao, Zhaoxiang Zhang, He Guan +2
Most of convolutional neural networks share the same characteristic: each convolutional layer is followed by a nonlinear activation layer where Rectified Linear Unit (ReLU) is the…
Poisoning Attacks to Local Differential Privacy for Ranking Estimation
Pei Zhan, Peng Tang, Yangzhuo Li +2
Local differential privacy (LDP) involves users perturbing their inputs to provide plausible deniability of their data. However, this also makes LDP vulnerable to poisoning attacks…
FedTopo: Topology-Informed Representation Alignment in Federated Learning under Non-I.I.D. Conditions
Ke Hu, Liyao Xiang, Peng Tang +1
Current federated-learning models deteriorate under heterogeneous (non-I.I.D.) client data, as their feature representations diverge and pixel- or patch-level objectives fail to ca…
Semi-Supervised Multi-Organ Segmentation via Deep Multi-Planar Co-Training
Yuyin Zhou, Yan Wang, Peng Tang +4
In multi-organ segmentation of abdominal CT scans, most existing fully supervised deep learning algorithms require lots of voxel-wise annotations, which are usually difficult, expe…
A Comprehensive Study on GDPR-Oriented Analysis of Privacy Policies: Taxonomy, Corpus and GDPR Concept Classifiers
Peng Tang, Xin Li, Yuxin Chen +5
Machine learning based classifiers that take a privacy policy as the input and predict relevant concepts are useful in different applications such as (semi-)automated compliance an…
OpenDerisk: An Industrial Framework for AI-Driven SRE, with Design, Implementation, and Case Studies
Peng Di, Faqiang Chen, Xiao Bai +12
The escalating complexity of modern software imposes an unsustainable operational burden on Site Reliability Engineering (SRE) teams, demanding AI-driven automation that can emulat…
An Adaptive Latent Factorization of Tensors Model for Embedding Dynamic Communication Network
Xin Liao, Qicong Hu, Peng Tang
The Dynamic Communication Network (DCN) describes the interactions over time among various communication nodes, and it is widely used in Big-data applications as a data source. As…
MoE-Prism: Disentangling Monolithic Experts for Elastic MoE Services via Model-System Co-Designs
Xinfeng Xia, Jiacheng Liu, Xiaofeng Hou +4
Mixture-of-Experts (MoE) models, the state-of-the-art in large-scale AI, achieve high quality by sparsely activating parameters. However, their reliance on routing between a few mo…
PCL: Proposal Cluster Learning for Weakly Supervised Object Detection
Peng Tang, Xinggang Wang, Song Bai +4
Weakly Supervised Object Detection (WSOD), using only image-level annotations to train object detectors, is of growing importance in object recognition. In this paper, we propose a…
Single-Shared Network with Prior-Inspired Loss for Parameter-Efficient Multi-Modal Imaging Skin Lesion Classification
Peng Tang, Tobias Lasser
In this study, we introduce a multi-modal approach that efficiently integrates multi-scale clinical and dermoscopy features within a single network, thereby substantially reducing…
Towards Personalized Differentially Private Learning for Decentralized Local Graphs
Longzhu He, Peng Tang, Chaozhuo Li +5
Graph-structured data is increasingly generated and stored in decentralized environments, such as social platforms, mobile applications, and edge networks, where users maintain con…
Towards Generalized Multi-Image Editing for Unified Multimodal Models
Pengcheng Xu, Peng Tang, Donghao Luo +7
Unified Multimodal Models (UMMs) integrate multimodal understanding and generation, yet they are limited to maintaining visual consistency and disambiguating visual cues when refer…
DocFormerv2: Local Features for Document Understanding
Srikar Appalaraju, Peng Tang, Qi Dong +3
We propose DocFormerv2, a multi-modal transformer for Visual Document Understanding (VDU). The VDU domain entails understanding documents (beyond mere OCR predictions) e.g., extrac…
FFP-300K: Scaling First-Frame Propagation for Generalizable Video Editing
Xijie Huang, Chengming Xu, Donghao Luo +6
First-Frame Propagation (FFP) offers a promising paradigm for controllable video editing, but existing methods are hampered by a reliance on cumbersome run-time guidance. We identi…
Proposal Learning for Semi-Supervised Object Detection
Peng Tang, Chetan Ramaiah, Yan Wang +2
In this paper, we focus on semi-supervised object detection to boost performance of proposal-based object detectors (a.k.a. two-stage object detectors) by training on both labeled…
DocKD: Knowledge Distillation from LLMs for Open-World Document Understanding Models
Sungnyun Kim, Haofu Liao, Srikar Appalaraju +6
Visual document understanding (VDU) is a challenging task that involves understanding documents across various modalities (text and image) and layouts (forms, tables, etc.). This s…
Object Detection in Videos by High Quality Object Linking
Peng Tang, Chunyu Wang, Xinggang Wang +3
Compared with object detection in static images, object detection in videos is more challenging due to degraded image qualities. An effective way to address this problem is to expl…
Self-locking non-volatile coding metasurfaces via origami-based mechanical bits
Ding Zhang, Peng Tang, Liqiao Jing +7
Digital coding metasurfaces have revolutionized electromagnetic (EM) manipulation, yet typical tunable approaches based on active components suffer from the "volatility" bottleneck…
Revisiting Multiple Instance Neural Networks
Xinggang Wang, Yongluan Yan, Peng Tang +2
Recently neural networks and multiple instance learning are both attractive topics in Artificial Intelligence related research fields. Deep neural networks have achieved great succ…
Feature Norm Regularized Federated Learning: Transforming Skewed Distributions into Global Insights
Ke Hu, WeiDong Qiu, Peng Tang
In the field of federated learning, addressing non-independent and identically distributed (non-i.i.d.) data remains a quintessential challenge for improving global model performan…
Robust Estimation of Sparse Precision Matrix using Adaptive Weighted Graphical Lasso Approach
Peng Tang, Huijing Jiang, Heeyoung Kim +1
Estimation of a precision matrix (i.e., inverse covariance matrix) is widely used to exploit conditional independence among continuous variables. The influence of abnormal observat…
PCR: A Prefetch-Enhanced Cache Reuse System for Low-Latency RAG Serving
Wenfeng Wang, Xiaofeng Hou, Peng Tang +5
Retrieval-Augmented Generation (RAG) systems enhance the performance of large language models (LLMs) by incorporating supplementary retrieved documents, enabling more accurate and…
DEED: Dynamic Early Exit on Decoder for Accelerating Encoder-Decoder Transformer Models
Peng Tang, Pengkai Zhu, Tian Li +3
Encoder-decoder transformer models have achieved great success on various vision-language (VL) tasks, but they suffer from high inference latency. Typically, the decoder takes up m…
Pay Less On Clinical Images: Asymmetric Multi-Modal Fusion Method For Efficient Multi-Label Skin Lesion Classification
Peng Tang, Tobias Lasser
Existing multi-modal approaches primarily focus on enhancing multi-label skin lesion classification performance through advanced fusion modules, often neglecting the associated ris…
LLMZero: Discovering Adaptive Training Strategies for RL Post-Training via LLM Agents
Haoyang Fang, Wei Zhu, Boran Han +11
RL post-training strategies are dataset-dependent and reveal a recurring empirical pattern: capacity parameters accumulate monotonically across stages, while regularization paramet…
MDS-GNN: A Mutual Dual-Stream Graph Neural Network on Graphs with Incomplete Features and Structure
Peng Yuan, Peng Tang
Graph Neural Networks (GNNs) have emerged as powerful tools for analyzing and learning representations from graph-structured data. A crucial prerequisite for the outstanding perfor…
Multiple Instance Detection Network with Online Instance Classifier Refinement
Peng Tang, Xinggang Wang, Xiang Bai +1
Of late, weakly supervised object detection is with great importance in object recognition. Based on deep learning, weakly supervised detectors have achieved many promising results…
Everyone's Privacy Matters! An Analysis of Privacy Leakage from Real-World Facial Images on Twitter and Associated User Behaviors
Yuqi Niu, Weidong Qiu, Peng Tang +5
Online users often post facial images of themselves and other people on online social networks (OSNs) and other Web 2.0 platforms, which can lead to potential privacy leakage of pe…
Open-Vocabulary Semantic Segmentation Network Integrating Object-Level Label and Scene-Level Semantic Features for Multimodal Remote Sensing Images
Jinkun Dai, Yuanxin Ye, Peng Tang +4
Semantic segmentation of multi-modal remote sensing imagery plays a pivotal role in land use/land cover (LULC) mapping, environmental monitoring, and precision earth observation. C…
FedDEAP: Adaptive Dual-Prompt Tuning for Multi-Domain Federated Learning
Yubin Zheng, Pak-Hei Yeung, Jing Xia +4
Federated learning (FL) enables multiple clients to collaboratively train machine learning models without exposing local data, balancing performance and privacy. However, domain sh…
The devil is in the details: Enhancing Video Virtual Try-On via Keyframe-Driven Details Injection
Qingdong He, Xueqin Chen, Yanjie Pan +7
Although diffusion transformer (DiT)-based video virtual try-on (VVT) has made significant progress in synthesizing realistic videos, existing methods still struggle to capture fin…
Automatic Fine-grained Glomerular Lesion Recognition in Kidney Pathology
Yang Nan, Fengyi Li, Peng Tang +5
Recognition of glomeruli lesions is the key for diagnosis and treatment planning in kidney pathology; however, the coexisting glomerular structures such as mesangial regions exacer…
Robustness of Object Recognition under Extreme Occlusion in Humans and Computational Models
Hongru Zhu, Peng Tang, Jeongho Park +2
Most objects in the visual world are partially occluded, but humans can recognize them without difficulty. However, it remains unknown whether object recognition models like convol…
Multi-Head Self-Attending Neural Tucker Factorization
Yikai Hou, Peng Tang
Quality-of-service (QoS) data exhibit dynamic temporal patterns that are crucial for accurately predicting missing values. These patterns arise from the evolving interactions betwe…
Shape-Texture Debiased Neural Network Training
Yingwei Li, Qihang Yu, Mingxing Tan +5
Shape and texture are two prominent and complementary cues for recognizing objects. Nonetheless, Convolutional Neural Networks are often biased towards either texture or shape, dep…
SE#PCFG: Semantically Enhanced PCFG for Password Analysis and Cracking
Yangde Wang, Weidong Qiu, Peng Tang +2
Much research has been done on user-generated textual passwords. Surprisingly, semantic information in such passwords remain under-investigated, with passwords created by English-…
Neural Canonical Polyadic Factorization for Traffic Analysis
Wenyu Luo, Yikai Hou, Peng Tang
Modern intelligent transportation systems rely on accurate spatiotemporal traffic analysis to optimize urban mobility and infrastructure resilience. However, pervasive missing data…
Deep FisherNet for Object Classification
Peng Tang, Xinggang Wang, Baoguang Shi +3
Despite the great success of convolutional neural networks (CNN) for the image classification task on datasets like Cifar and ImageNet, CNN's representation power is still somewhat…