1 citations · 1 across the 2 of their papers we have counts for
5 papers
LayoutBench: Performance Benchmarking of Cloud Storage Layouts for Multimedia Data
Debopam Sanyal, Hongjie Chen, Alexey Tumanov +1
Modern multimedia machine learning workloads increasingly store large-scale datasets in cloud object storage services such as AWS S3. How these samples are physically organized in…
Rethinking Dataset Distillation for Classification: Do Distilled Sets Outperform Coresets?
Trisha Mittal, Akshay Mehra, Joshua Kimball
Dataset distillation (DD) has emerged as a prominent approach in data centric machine learning, aiming to synthesize compact training sets for efficient training by compressing the…
Variable-Length Audio Fingerprinting
Hongjie Chen, Hanyu Meng, Huimin Zeng +3
Audio fingerprinting converts audio to much lower-dimensional representations, allowing distorted recordings to still be recognized as their originals through similar fingerprints.…
Measuring Time-Series Dataset Similarity using Wasserstein Distance
Hongjie Chen, Akshay Mehra, Josh Kimball +1
The emergence of time-series foundation model research elevates the growing need to measure the (dis)similarity of time-series datasets. A time-series dataset similarity measure ai…
Coreset Selection via LLM-based Concept Bottlenecks
Akshay Mehra, Trisha Mittal, Subhadra Gopalakrishnan +1
Coreset Selection (CS) aims to identify a subset of the training dataset that achieves model performance comparable to using the entire dataset. Many state-of-the-art CS methods se…