6 papers
Decomposing the Depth Profile of Fine-Tuning
Jayadev Billa
Fine-tuning adapts pretrained networks to new objectives. Whether the resulting depth profile of representational change reflects an intrinsic property of the model or the magnitud…
Predicting Where Steering Vectors Succeed
Jayadev Billa
Steering vectors work for some concepts and layers but fail for others, and practitioners have no way to predict which setting applies before running an intervention. We introduce…
The Geometric Anatomy of Capability Acquisition in Transformers
Jayadev Billa
Neural networks gain capabilities during training, but the internal changes that precede capability acquisition are not well understood. In particular, the relationship between geo…
When Audio-LLMs Don't Listen: A Cross-Linguistic Study of Modality Arbitration
Jayadev Billa
When audio and text conflict, speech-enabled language models follow text far more often than they do when arbitrating between two conflicting text sources, even under explicit inst…
Modality Collapse as Mismatched Decoding: Information-Theoretic Limits of Multimodal LLMs
Jayadev Billa
Numerous studies have shown that multimodal LLMs process speech and images well but fail in non-intuitive ways rendering trivial tasks such as object counting unreliable. We invest…
The Cascade Equivalence Hypothesis: When Do Speech LLMs Behave Like ASRLLM Pipelines?
Jayadev Billa
Speech LLMs are widely understood to be better than ASRLLM cascades since they have access to the audio directly, and not just the transcript. In this paper, we presen…