5 papers
Language-Aware Distillation for Multilingual Instruction-Following Speech LLMs with ASR-Only Supervision
Shreyas Gopal, Donghang Wu, Ashutosh Anshul +5
Speech Large Language Models (LLMs) that understand and follow instructions in many languages are useful for real-world interaction, but are difficult to train with supervised fine…
Next-Frame Feature Prediction for Multimodal Deepfake Detection and Temporal Localization
Ashutosh Anshul, Shreyas Gopal, Deepu Rajan +1
Recent multimodal deepfake detection methods designed for generalization conjecture that single-stage supervised training struggles to generalize across unseen manipulations and da…
A Multimodal Framework for Depression Detection during Covid-19 via Harvesting Social Media: A Novel Dataset and Method
Ashutosh Anshul, Gumpili Sai Pranav, Mohammad Zia Ur Rehman +1
The recent coronavirus disease (Covid-19) has become a pandemic and has affected the entire globe. During the pandemic, we have observed a spike in cases related to mental health,…
Explainable Disentanglement on Discrete Speech Representations for Noise-Robust ASR
Shreyas Gopal, Ashutosh Anshul, Haoyang Li +3
Discrete audio representations are gaining traction in speech modeling due to their interpretability and compatibility with large language models, but are not always optimized for…
RoGBot: Relationship-Oblivious Graph-based Neural Network with Contextual Knowledge for Bot Detection
Ashutosh Anshul, Mohammad Zia Ur Rehman, Sri Akash Kadali +1
Detecting automated accounts (bots) among genuine users on platforms like Twitter remains a challenging task due to the evolving behaviors and adaptive strategies of such accounts.…