1 citations · 2 across the 5 of their papers we have counts for
5 papers
Diffusion-Augmented Coreset Expansion for Scalable Dataset Distillation
Ali Abbasi, Shima Imani, Chenyang An +6
With the rapid scaling of neural networks, data storage and communication demands have intensified. Dataset distillation has emerged as a promising solution, condensing information…
Next-Token Prediction Task Assumes Optimal Data Ordering for LLM Training in Proof Generation
Chenyang An, Shima Imani, Feng Yao +8
In the field of large language model (LLM)-based proof generation, despite extensive training on large datasets such as ArXiv, LLMs still exhibit only modest performance on proving…
Learning How To Ask: Cycle-Consistency Refines Prompts in Multimodal Foundation Models
Maurice Diesendruck, Jianzhe Lin, Shima Imani +3
When LLMs perform zero-shot inference, they typically use a prompt with a task specification, and generate a completion. However, there is no work to explore the possibility of the…
BatchPrompt: Accomplish more with less
Jianzhe Lin, Maurice Diesendruck, Liang Du +1
As the ever-increasing token limits of large language models (LLMs) have enabled long context as input, prompting with single data samples might no longer an efficient way. A strai…
Discovering Distribution Shifts using Latent Space Representations
Leo Betthauser, Urszula Chajewska, Maurice Diesendruck +1
Rapid progress in representation learning has led to a proliferation of embedding models, and to associated challenges of model selection and practical application. It is non-trivi…