papers

Publications (21)

cs.LG2024

Residual vector quantization for KV cache compression in large language model

Ankur Kumar

KV cache compression methods have mainly relied on scalar quantization techniques to reduce the memory requirements during decoding. In this work, we apply residual vector quantiza…

cs.PL2022

LabVIEW is faster and C is economical interfacing tool for UCT automation

Ankur Kumar, Mayank Goswami

An in-house developed 2D ultrasound computerized Tomography system is fully automated. Performance analysis of instrument and software interfacing soft tools, namely the LabVIEW, M…

cs.LG2021

A review of on-device fully neural end-to-end automatic speech recognition algorithms

Chanwoo Kim, Dhananjaya Gowda, Dongsoo Lee +5

In this paper, we review various end-to-end automatic speech recognition algorithms and their optimization techniques for on-device applications. Conventional speech recognition sy…

astro-ph.EP2026

Towards a Foundation Model for the Martian Atmosphere

Sujit Roy, Udayshankar Nair, Yuling Wu +16

The martian atmosphere hosts dynamical phenomena ranging from planet-encircling dust storms to mesoscale orographic clouds and nocturnal low-level jets. General circulation model s…

eess.AS2025

A Data-Driven Diffusion-based Approach for Audio Deepfake Explanations

Petr Grinberg, Ankur Kumar, Surya Koppisetti +1

Evaluating explainability techniques, such as SHAP and LRP, in the context of audio deepfake detection is challenging due to lack of clear ground truth annotations. In the cases wh…

physics.flu-dyn2025

Experimental Investigation of Acoustically Forced Helium Jet in Crossflow Using Shadowgraphy and Modal Analysis

Ankur Kumar, Sita Ram Sahu, Narsing K Jha +1

This study presents an experimental investigation of helium jet in crossflow of air, with an objective to understand the influence of acoustic forcing on jet behavior and mixing. U…

cs.CV2025

Multi Attribute Bias Mitigation via Representation Learning

Rajeev Ranjan Dwivedi, Ankur Kumar, Vinod K Kurmi

Real world images frequently exhibit multiple overlapping biases, including textures, watermarks, gendered makeup, scene object pairings, etc. These biases collectively impair the…

physics.flu-dyn2024

Reacting Hydrogen Jet in Crossflow - Flame Dynamics Under Acoustic Forcing

Ankur Kumar, Anubhav sinha

This paper presents experimental study of reacting hydrogen jet in crossflow. High speed shadowgraph images are used to capture flame dynamics. Unforced jets with various momentum…

eess.AS2019

Improved Multi-Stage Training of Online Attention-based Encoder-Decoder Models

Abhinav Garg, Dhananjaya Gowda, Ankur Kumar +3

In this paper, we propose a refined multi-stage multi-task training strategy to improve the performance of online attention-based encoder-decoder (AED) models. A three-stage traini…

cs.CL2024

Stateful Conformer with Cache-based Inference for Streaming Automatic Speech Recognition

Vahid Noroozi, Somshubra Majumdar, Ankur Kumar +2

In this paper, we propose an efficient and accurate streaming speech recognition model based on the FastConformer architecture. We adapted the FastConformer architecture for stream…

cs.LG2025

AI Agents in Drug Discovery

Srijit Seal, Dinh Long Huynh, Moudather Chelbi +17

Artificial intelligence (AI) agents are emerging as transformative tools in drug discovery, with the ability to autonomously reason, act, and learn through complicated research wor…

q-bio.QM2024

Cell Painting Gallery: an open resource for image-based profiling

Erin Weisbart, Ankur Kumar, John Arevalo +3

Image-based or morphological profiling is a rapidly expanding field wherein cells are "profiled" by extracting hundreds to thousands of unbiased, quantitative features from images…

eess.SP2021

AI and conventional methods for UCT projection data estimation

Ankur Kumar, Prasunika Khare, Mayank Goswami

A 2D Compact ultrasound computerized tomography (UCT) system is developed. Fully automatic post processing tools involving signal and image processing are developed as well. Square…

cs.LG2025

What Does an Audio Deepfake Detector Focus on? A Study in the Time Domain

Petr Grinberg, Ankur Kumar, Surya Koppisetti +1

Adding explanations to audio deepfake detection (ADD) models will boost their real-world application by providing insight on the decision making process. In this paper, we propose…

cs.CV2025

Contrasting Low and High-Resolution Features for HER2 Scoring using Deep Learning

Ekansh Chauhan, Anila Sharma, Amit Sharma +7

Breast cancer, the most common malignancy among women, requires precise detection and classification for effective treatment. Immunohistochemistry (IHC) biomarkers like HER2, ER, a…

physics.med-ph2024

Pulse excitation mode selection via AI Pipeline to Fully Automate the WUCT System

Ankur Kumar, Mayank Goswami

The parametric optimization for the ultrasound computed tomography system is introduced. It is hypothesized that the pulse characteristic directly affects the information present i…

cs.CV2026

Prithvi-EO-2.0: A Versatile Multi-Temporal Foundation Model for Earth Observation Applications

Daniela Szwarcman, Sujit Roy, Paolo Fraccaro +33

This paper presents Prithvi-EO-2.0, a new geospatial foundation model that offers significant improvements over its predecessor, Prithvi-EO-1.0. Trained on 4.2 million global time…

cs.IR2026

Scientific Code Search at Scale: A Multi-Domain Dataset and Benchmark

Nishan Pantha, Pranath Reddy Kumbam, Sajil Awale +8

Scientists increasingly rely on open-source tools to support their research workflows, yet discovering relevant software among over 600 million GitHub repositories remains challeng…

cs.CV2022

Vision Transformer Compression with Structured Pruning and Low Rank Approximation

Ankur Kumar

Transformer architecture has gained popularity due to its ability to scale with large dataset. Consequently, there is a need to reduce the model size and latency, especially for on…

eess.AS2023

Fast Conformer with Linearly Scalable Attention for Efficient Speech Recognition

Dima Rekesh, Nithin Rao Koluguri, Samuel Kriman +8

Conformer-based models have become the dominant end-to-end architecture for speech processing tasks. With the objective of enhancing the conformer architecture for efficient traini…

eess.AS2024

Learn from Real: Reality Defender's Submission to ASVspoof5 Challenge

Yi Zhu, Chirag Goel, Surya Koppisetti +3

Audio deepfake detection is crucial to combat the malicious use of AI-synthesized speech. Among many efforts undertaken by the community, the ASVspoof challenge has become one of t…