2 papers
cs.SE2026
BlueFin: Benchmarking LLM Agents on Financial Spreadsheets
Srivatsa Kundurthy, Clara Na, Colton Moraine +6
We present BlueFin, a benchmark that tasks large language model (LLM) agents with synthesis, manipulation, and comprehension tasks over spreadsheet workbooks in the professional fi…
cs.LG2025
Attribution-Guided Distillation of Matryoshka Sparse Autoencoders
Cristina P. Martin-Linares, Jonathan P. Ling
Sparse autoencoders (SAEs) aim to disentangle model activations into monosemantic, human-interpretable features. In practice, learned features are often redundant and vary across t…