activity
20232026
most citedDiff-SAGe: End-to-End Spatial Audio Generation Using Diffusion Models

4 citations · 4 across the 6 of their papers we have counts for

collaborators
Showing cs.SDShow all

5 papers · 1 filter

cs.SD2026

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text

Weiguo Pian, Saksham Singh Kushwaha, Zhimin Chen +4

In this paper, we propose Universal Holistic Audio Generation (UniHAGen), a task for synthesizing comprehensive auditory scenes that include both on-screen and off-screen sounds ac…

cs.SD2025

: Benchmarking the Robustness of Audio-Visual Recognition Models at Test-Time

Sarthak Kumar Maharana, Saksham Singh Kushwaha, Baoming Zhang +4

While recent audio-visual models have demonstrated impressive performance, their robustness to distributional shifts at test-time remains not fully understood. Existing robustness…

cs.SD2024★ 4 cited

Diff-SAGe: End-to-End Spatial Audio Generation Using Diffusion Models

Saksham Singh Kushwaha, Jianbo Ma, Mark R. P. Thomas +2

Spatial audio is a crucial component in creating immersive experiences. Traditional simulation-based approaches to generate spatial audio rely on expertise, have limited scalabilit…

cs.SD2023

Sound Source Distance Estimation in Diverse and Dynamic Acoustic Conditions

Saksham Singh Kushwaha, Iran R. Roman, Magdalena Fuentes +1

Localizing a moving sound source in the real world involves determining its direction-of-arrival (DOA) and distance relative to a microphone. Advancements in DOA estimation have be…

cs.SD2023

A Multimodal Prototypical Approach for Unsupervised Sound Classification

Saksham Singh Kushwaha, Magdalena Fuentes

In the context of environmental sound classification, the adaptability of systems is key: which sound classes are interesting depends on the context and the user's needs. Recent ad…