8 papers
USAD: Uncertainty-aware Statistical Adversarial Detection
Zhijian Zhou, Xunye Tian, Jiacheng Zhang +5
Statistical adversarial detection (SAD) treats detection as a two-sample test. Given a reference set of clean examples (CEs) and a batch of queries, potentially containing an unkno…
CELEUS: Certifiable and Efficient LLM Evaluation via E-Processes
Zhijian Zhou, Zesheng Ye, Zhaorun Chen +2
Can we trust evaluation scores to capture an LLM's true real-world performance? Certifiable evaluation answers this question by providing guarantee for LLM evaluation. In particula…
What Do Deepfake Benchmarks Measure? An Audit Using Frozen Self-Supervised Representations
Samuel Pagon, Yixuan Shen, Vishal Asnani +1
As deepfake generators approach perceptual indistinguishability, reliable detection becomes critical. Yet, detectors that score well on benchmarks routinely fail in the wild. A con…
FedReLa: Imbalanced Federated Learning via Re-Labeling
Guangzheng Hu, Patricia Menéndez, Feng Liu +3
Federated learning has emerged as the foremost approach for decentralized model training with privacy preservation. The global class imbalance and cross-client data heterogeneity n…
Semantic Robustness Certification for Vision-Language Models
Peiyu Yang, Paul Montague, Feng Liu +4
Vision-language models (VLMs) are now widely used in downstream tasks. However, real-world applications often expose VLMs to distribution shifts induced by semantic variation (e.g.…
Are Two Datasets Close Enough With Statistical Significance? A Kernel Distributional Closeness Testing Approach
Zhijian Zhou, Liuhua Peng, Xunye Tian +2
Are two distributions close to each other with statistical significance? Distribution closeness testing (DCT) formalizes this question by testing whether the distance between a dis…