7 papers
ICLAD: In-Context Learning with Comparison-Guidance for Audio Deepfake Detection
Benjamin Chou, Yi Zhu, Surya Koppisetti
Audio deepfakes pose a significant security threat, yet current state-of-the-art (SOTA) detection systems do not generalize well to realistic in-the-wild deepfakes. We introduce a…
Advancing Multi-Instrument Music Transcription: Results from the 2025 AMT Challenge
Ojas Chaturvedi, Kayshav Bhardwaj, Tanay Gondil +5
This paper presents the results of the 2025 Automatic Music Transcription (AMT) Challenge, an online competition to benchmark progress in multi-instrument transcription. Eight team…
LadderSym: A Multimodal Interleaved Transformer for Music Practice Error Detection
Benjamin Shiue-Hal Chou, Purvish Jajal, Nick John Eliopoulos +4
Music learners can greatly benefit from tools that accurately detect errors in their practice. Existing approaches typically compare audio recordings to music scores using heuristi…
Inference-Time Alignment of Diffusion Models via Evolutionary Algorithms
Purvish Jajal, Nick John Eliopoulos, Benjamin Shiue-Hal Chou +3
Diffusion models are state-of-the-art generative models, yet their samples often fail to satisfy application objectives such as safety constraints or domain-specific validity. Exis…
AdaPerceiver: Transformers with Adaptive Width, Depth, and Tokens
Purvish Jajal, Nick John Eliopoulos, Benjamin Shiue-Hal Chou +3
Modern transformer architectures achieve remarkable performance across tasks and domains but remain rigid in how they allocate computation at inference time. Real-world deployment…
Token Turing Machines are Efficient Vision Models
Purvish Jajal, Nick John Eliopoulos, Benjamin Shiue-Hal Chou +3
We propose Vision Token Turing Machines (ViTTM), an efficient, low-latency, memory-augmented Vision Transformer (ViT). Our approach builds on Neural Turing Machines and Token Turin…