22 papers
Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks
Nobin Sarwar, Shubhashis Roy Dipta, Zheyuan Liu +1
With the growing adoption of VLMs, DMs, LLMs, and AFMs, these multimodal foundation models can inadvertently encode sensitive, copyrighted, biased, or unsafe cross-modal associatio…
SpreadsheetArena: Decomposing Preference in LLM Generation of Spreadsheet Workbooks
Srivatsa Kundurthy, Clara Na, Michael Handley +5
We consider the task of end-to-end spreadsheet generation, where language models produce spreadsheet artifacts to satisfy users' explicit and implicit constraints, specified in nat…
Task Decomposition for Efficient Annotation
Nupoor Gandhi, Emma Strubell
High-quality annotations of structured representations are expensive to collect over large corpora. Manual annotation of structure is laborious, and model-based annotation, althoug…
Evaluation of ML Resource Utilization Requires Model Life Cycle Assessment
Jared Fernandez, Clara Na, Yonatan Bisk +2
Proper accounting of the energy requirements and environmental impact of artificial intelligence (AI) systems is necessary for researchers, developers, policy makers, and users to…
BlueFin: Benchmarking LLM Agents on Financial Spreadsheets
Srivatsa Kundurthy, Clara Na, Colton Moraine +6
We present BlueFin, a benchmark that tasks large language model (LLM) agents with synthesis, manipulation, and comprehension tasks over spreadsheet workbooks in the professional fi…
The Hidden Cost of Thinking: Energy Use and Environmental Impact of LMs Beyond Pretraining
Jacob Morrison, Noah A. Smith, Emma Strubell
Modern language model development extends far beyond pretraining, yet environmental reporting remains narrowly focused on the cost of training a single final model. In this work, w…