papers

Publications (6)

cs.SD2023

FlexiAST: Flexibility is What AST Needs

Jiu Feng, Mehmet Hamza Erol, Joon Son Chung +1

The objective of this work is to give patch-size flexibility to Audio Spectrogram Transformers (AST). Recent advancements in ASTs have shown superior performance in various audio-b…

cs.CV2026

Egocentric World Model for Photorealistic Hand-Object Interaction Synthesis

Dayou Li, Lulin Liu, Bangya Liu +6

To serve as a scalable data source for embodied AI, world models should act as true simulators that infer interaction dynamics strictly from user actions, rather than mere conditio…

cs.SD2024

ElasticAST: An Audio Spectrogram Transformer for All Length and Resolutions

Jiu Feng, Mehmet Hamza Erol, Joon Son Chung +1

Transformers have rapidly overtaken CNN-based architectures as the new standard in audio classification. Transformer-based models, such as the Audio Spectrogram Transformers (AST),…

cs.SD2024

From Coarse to Fine: Efficient Training for Audio Spectrogram Transformers

Jiu Feng, Mehmet Hamza Erol, Joon Son Chung +1

Transformers have become central to recent advances in audio classification. However, training an audio spectrogram transformer, e.g. AST, from scratch can be resource and time-int…

cs.SD2024

Audio Mamba: Bidirectional State Space Model for Audio Representation Learning

Mehmet Hamza Erol, Arda Senocak, Jiu Feng +1

Transformers have rapidly become the preferred choice for audio classification, surpassing methods based on CNNs. However, Audio Spectrogram Transformers (ASTs) exhibit quadratic s…

cs.CV2022

Decoupled Adversarial Contrastive Learning for Self-supervised Adversarial Robustness

Chaoning Zhang, Kang Zhang, Chenshuang Zhang +4

Adversarial training (AT) for robust representation learning and self-supervised learning (SSL) for unsupervised representation learning are two active research fields. Integrating…