activity
20242026
collaborators

9 papers

cs.CR2026

Taming Data Challenges in ML-based Security Tasks Using Generative AI

Shravya Kanchi, Neal Mangaokar, Aravind Cheruvu +4

Machine learning-based supervised classifiers are widely used for security tasks, and their improvement has been largely focused on algorithmic advancements. We argue that data cha…

cs.CV2025

Test-Time Canonicalization by Foundation Models for Robust Perception

Utkarsh Singhal, Ryan Feng, Stella X. Yu +1

Perception in the real world requires robustness to diverse viewing conditions. Existing approaches often rely on specialized architectures or training with predefined data augment…

cs.LG2025

Class-Proportional Coreset Selection for Difficulty-Separable Data

Elisa Tsai, Haizhong Zheng, Atul Prakash

High-quality training data is essential for building reliable and efficient machine learning systems. One-shot coreset selection addresses this by pruning the dataset while maintai…

cs.LG2025

What Really is a Member? Discrediting Membership Inference via Poisoning

Neal Mangaokar, Ashish Hooda, Zhuohang Li +5

Membership inference tests aim to determine whether a particular data point was included in a language model's training set. However, recent works have shown that such tests often…

cs.CL2025

Plato: Plan to Efficiently Decode for Large Language Model Inference

Shuowei Jin, Xueshen Liu, Yongji Wu +7

Large language models (LLMs) have achieved remarkable success in natural language tasks, but their inference incurs substantial computational and memory overhead. To improve effici…

cs.CV2025

ELFS: Label-Free Coreset Selection with Proxy Training Dynamics

Haizhong Zheng, Elisa Tsai, Yifu Lu +4

High-quality human-annotated data is crucial for modern deep learning pipelines, yet the human annotation process is both costly and time-consuming. Given a constrained human label…