9 papers
Taming Data Challenges in ML-based Security Tasks Using Generative AI
Shravya Kanchi, Neal Mangaokar, Aravind Cheruvu +4
Machine learning-based supervised classifiers are widely used for security tasks, and their improvement has been largely focused on algorithmic advancements. We argue that data cha…
Test-Time Canonicalization by Foundation Models for Robust Perception
Utkarsh Singhal, Ryan Feng, Stella X. Yu +1
Perception in the real world requires robustness to diverse viewing conditions. Existing approaches often rely on specialized architectures or training with predefined data augment…
Class-Proportional Coreset Selection for Difficulty-Separable Data
Elisa Tsai, Haizhong Zheng, Atul Prakash
High-quality training data is essential for building reliable and efficient machine learning systems. One-shot coreset selection addresses this by pruning the dataset while maintai…
What Really is a Member? Discrediting Membership Inference via Poisoning
Neal Mangaokar, Ashish Hooda, Zhuohang Li +5
Membership inference tests aim to determine whether a particular data point was included in a language model's training set. However, recent works have shown that such tests often…
Plato: Plan to Efficiently Decode for Large Language Model Inference
Shuowei Jin, Xueshen Liu, Yongji Wu +7
Large language models (LLMs) have achieved remarkable success in natural language tasks, but their inference incurs substantial computational and memory overhead. To improve effici…
ELFS: Label-Free Coreset Selection with Proxy Training Dynamics
Haizhong Zheng, Elisa Tsai, Yifu Lu +4
High-quality human-annotated data is crucial for modern deep learning pipelines, yet the human annotation process is both costly and time-consuming. Given a constrained human label…