most citedWhisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music

2 citations · 2 across the 1 of their papers we have counts for

collaborators

7 papers

cs.SD20262 cited

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music

Prateek Verma

We propose WHISPER-GPT: A generative large language model (LLM) for speech and music that allows us to work with continuous audio representations and discrete tokens simultaneously…

cs.LG2025

Fine-Tuning Vision-Language Models for Multimodal Polymer Property Prediction

An Vuong, Minh-Hao Van, Prateek Verma +2

Vision-Language Models (VLMs) have shown strong performance in tasks like visual question answering and multimodal text generation, but their effectiveness in scientific domains su…

cs.SD2025

Thinking While Listening: Simple Test Time Scaling For Audio Classification

Prateek Verma, Mert Pilanci

We propose a framework that enables neural models to "think while listening" to everyday sounds, thereby enhancing audio classification performance. Motivated by recent advances in…

cs.CL2025

Large Language Models Implicitly Learn to See and Hear Just By Reading

Prateek Verma, Mert Pilanci

This paper presents a fascinating find: By training an auto-regressive LLM model on text tokens, the text model inherently develops internally an ability to understand images and a…

cs.LG2025

A Survey of AI for Materials Science: Foundation Models, LLM Agents, Datasets, and Tools

Minh-Hao Van, Prateek Verma, Chen Zhao +1

Foundation models (FMs) are catalyzing a transformative shift in materials science (MatSci) by enabling scalable, general-purpose, and multimodal AI systems for scientific discover…

cs.IT2025

Semantic De-boosting in e-commerce Query Autocomplete

Adithya Rajan, Weiqi Tong, Greg Sharp +2

In ecommerce search, query autocomplete plays a critical role to help users in their shopping journey. Often times, query autocomplete presents users with semantically similar quer…