2 citations · 2 across the 1 of their papers we have counts for
7 papers
Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music
Prateek Verma
We propose WHISPER-GPT: A generative large language model (LLM) for speech and music that allows us to work with continuous audio representations and discrete tokens simultaneously…
Fine-Tuning Vision-Language Models for Multimodal Polymer Property Prediction
An Vuong, Minh-Hao Van, Prateek Verma +2
Vision-Language Models (VLMs) have shown strong performance in tasks like visual question answering and multimodal text generation, but their effectiveness in scientific domains su…
Thinking While Listening: Simple Test Time Scaling For Audio Classification
Prateek Verma, Mert Pilanci
We propose a framework that enables neural models to "think while listening" to everyday sounds, thereby enhancing audio classification performance. Motivated by recent advances in…
Large Language Models Implicitly Learn to See and Hear Just By Reading
Prateek Verma, Mert Pilanci
This paper presents a fascinating find: By training an auto-regressive LLM model on text tokens, the text model inherently develops internally an ability to understand images and a…
A Survey of AI for Materials Science: Foundation Models, LLM Agents, Datasets, and Tools
Minh-Hao Van, Prateek Verma, Chen Zhao +1
Foundation models (FMs) are catalyzing a transformative shift in materials science (MatSci) by enabling scalable, general-purpose, and multimodal AI systems for scientific discover…
Semantic De-boosting in e-commerce Query Autocomplete
Adithya Rajan, Weiqi Tong, Greg Sharp +2
In ecommerce search, query autocomplete plays a critical role to help users in their shopping journey. Often times, query autocomplete presents users with semantically similar quer…