3 papers
cs.LG2025
Tracing the Representation Geometry of Language Models from Pretraining to Post-training
Melody Zixuan Li, Kumar Krishna Agrawal, Arna Ghosh +4
Standard training metrics like loss fail to explain the emergence of complex capabilities in large language models. We take a spectral approach to investigate the geometry of learn…
cs.LG2025
Harnessing small projectors and multiple views for efficient vision pretraining
Kumar Krishna Agrawal, Arna Ghosh, Shagun Sodhani +2
Recent progress in self-supervised (SSL) visual representation learning has led to the development of several different proposed frameworks that rely on augmentations of images but…
cs.LG2024
Attribute Diversity Determines the Systematicity Gap in VQA
Ian Berlot-Attwell, Kumar Krishna Agrawal, A. Michael Carrell +2
Although modern neural networks often generalize to new combinations of familiar concepts, the conditions that enable such compositionality have long been an open question. In this…