collaborators

6 papers

cs.LG2026

Decomposing the Depth Profile of Fine-Tuning

Jayadev Billa

Fine-tuning adapts pretrained networks to new objectives. Whether the resulting depth profile of representational change reflects an intrinsic property of the model or the magnitud…

cs.LG2026

Predicting Where Steering Vectors Succeed

Jayadev Billa

Steering vectors work for some concepts and layers but fail for others, and practitioners have no way to predict which setting applies before running an intervention. We introduce…

cs.LG2026

The Geometric Anatomy of Capability Acquisition in Transformers

Jayadev Billa

Neural networks gain capabilities during training, but the internal changes that precede capability acquisition are not well understood. In particular, the relationship between geo…

cs.CL2026

When Audio-LLMs Don't Listen: A Cross-Linguistic Study of Modality Arbitration

Jayadev Billa

When audio and text conflict, speech-enabled language models follow text far more often than they do when arbitrating between two conflicting text sources, even under explicit inst…

cs.CL2026

Modality Collapse as Mismatched Decoding: Information-Theoretic Limits of Multimodal LLMs

Jayadev Billa

Numerous studies have shown that multimodal LLMs process speech and images well but fail in non-intuitive ways rendering trivial tasks such as object counting unreliable. We invest…

cs.CL2026

The Cascade Equivalence Hypothesis: When Do Speech LLMs Behave Like ASRLLM Pipelines?

Jayadev Billa

Speech LLMs are widely understood to be better than ASRLLM cascades since they have access to the audio directly, and not just the transcript. In this paper, we presen…