4 papers
Merge++: Universal Merge Refinement Through Data-Free Checkpoint Inversion
Aditya Pola, Vineeth N. Balasubramanian
Model merging consolidates fine-tuned experts into one multi-task model without retraining. All existing data-free methods approach this problem entirely in weight space. Restricte…
ISA-Bench: A Benchmark for Computational Reasoning Across Instruction Set Architectures
Aditya Pola, Arkaprava Majumdar, Vineeth N. Balasubramanian
Large language model code generation benchmarks primarily evaluate well-resourced languages like Python and Java, where models benefit from abundant training data. They provide lim…
LogicCBMs: Logic-Enhanced Concept-Based Learning
Deepika SN Vemuri, Gautham Bellamkonda, Aditya Pola +1
Concept Bottleneck Models (CBMs) provide a basis for semantic abstractions within a neural network architecture. Such models have primarily been seen through the lens of interpreta…
Where does an LLM begin computing an instruction?
Aditya Pola, Vineeth N. Balasubramanian
Following an instruction involves distinct sub-processes, such as reading content, reading the instruction, executing it, and producing an answer. We ask where, along the layer sta…