activity
20242026
most citedMeta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound

1 citations · 1 across the 4 of their papers we have counts for

collaborators

6 papers

eess.AS2026

SAM Audio Judge: A Unified Multimodal Framework for Perceptual Evaluation of Audio Separation

Helin Wang, Bowen Shi, Andros Tjandra +6

The performance evaluation remains a complex challenge in audio separation, and existing evaluation metrics are often misaligned with human perception, course-grained, relying on g…

cs.LG2026

Mosaic: Unlocking Long-Context Inference for Diffusion LLMs via Global Memory Planning and Dynamic Peak Taming

Liang Zheng, Bowen Shi, Yitao Hu +5

Diffusion-based large language models (dLLMs) have emerged as a promising paradigm, utilizing simultaneous denoising to enable global planning and iterative refinement. While these…

eess.AS2025

SAM Audio: Segment Anything in Audio

Bowen Shi, Andros Tjandra, John Hoffman +11

General audio source separation is a key capability for multimodal AI systems that can perceive and reason about sound. Despite substantial progress in recent years, existing separ…

cs.SD2025

MR-FlowDPO: Multi-Reward Direct Preference Optimization for Flow-Matching Text-to-Music Generation

Alon Ziv, Sanyuan Chen, Andros Tjandra +3

A key challenge in music generation models is their lack of direct alignment with human preferences, as music evaluation is inherently subjective and varies widely across individua…

cs.SD20251 cited

Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound

Andros Tjandra, Yi-Chiao Wu, Baishan Guo +10

The quantification of audio aesthetics remains a complex challenge in audio processing, primarily due to its subjective nature, which is influenced by human perception and cultural…

eess.AS2024

Audiobox TTA-RAG: Improving Zero-Shot and Few-Shot Text-To-Audio with Retrieval-Augmented Generation

Mu Yang, Bowen Shi, Matthew Le +2

This work focuses on improving Text-To-Audio (TTA) generation on zero-shot and few-shot settings (i.e. generating unseen or uncommon audio events). Inspired by the success of Retri…