6 papers
Know Thy Reasoner: Not All Language Models Explore Alike
Moulik Choraria, Argyrios Gerogiannis, Anirban Das +4
Compute scaling for LLM reasoning trades off exploring solution approaches (\emph{breadth}) against refining promising ones (\emph{depth}), yet why a given trade-off works, and why…
Context-Gated Associative Retrieval: From Theory to Transformers
Moulik Choraria, Argyrios Gerogiannis, Vidhata Jayaraman +2
Hopfield networks and their generalizations have established deep connections among biological associative memories, statistical physics, and transformers. Yet most models treat re…
Skip-It? Theoretical Conditions for Layer Skipping in Vision-Language Models
Max Hartman, Vidhata Jayaraman, Moulik Choraria +2
Vision-language models achieve incredible performance across a wide range of tasks, but their large size makes inference costly. Recent work has shown that multimodal processing co…
Watermarking Discrete Diffusion Language Models
Avi Bagchi, Akhil Bhimaraju, Moulik Choraria +2
Watermarking has emerged as a promising technique to track AI-generated content and differentiate it from authentic human creations. While prior work extensively studies watermarki…
DeepInsert: Early Layer Bypass for Efficient and Performant Multimodal Understanding
Moulik Choraria, Xinbo Wu, Akhil Bhimaraju +5
Hyperscaling of data and parameter count in LLMs is yielding diminishing improvement when weighed against training costs, underlining a growing need for more efficient finetuning a…
Semantically Grounded QFormer for Efficient Vision Language Understanding
Moulik Choraria, Xinbo Wu, Sourya Basu +5
General purpose Vision Language Models (VLMs) have received tremendous interest in recent years, owing to their ability to learn rich vision-language correlations as well as their…