collaborators

13 papers

cs.AI2026

SemPlan: Benchmarking Structured Semantic Planning for LLM-Based Queries over Enterprise Data

Bruno Santos Teixeira

Natural-language interfaces to enterprise data must translate underspecified requests into governed, executable behavior while controlling invalid queries, policy failures, cost, a…

cs.SD2026

ALM2Vec: Learning Audio Embeddings for Universal Audio Retrieval with Large Audio-Language Models

Fengjie Lu, Chenang Jiang, Jiarui Hai +2

Recent advances in language--audio retrieval have been largely driven by contrastive dual-encoder architectures that align audio and text in a shared embedding space. While effecti…

cs.SD2026

AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking

Xilin Jiang, Qiaolin Wang, Junkai Wu +30

Internet audio-visual clips convey meaning through time-varying sound and motion, which extend beyond what text alone can represent. To examine whether AI models can understand suc…

cs.SD2026

Summary of The Inaugural Music Source Restoration Challenge

Yongyi Zang, Jiarui Hai, Wanying Ge +5

Music Source Restoration (MSR) aims to recover original, unprocessed instrument stems from professionally mixed and degraded audio, requiring the reversal of both production effect…

cs.SD2025

Adapting Speech Language Model to Singing Voice Synthesis

Yiwen Zhao, Jiatong Shi, Jinchuan Tian +4

Speech Language Models (SLMs) have recently emerged as a unified paradigm for addressing a wide range of speech-related tasks, including text-to-speech (TTS), speech enhancement (S…

cs.SD2025

MSRBench: A Benchmarking Dataset for Music Source Restoration

Yongyi Zang, Jiarui Hai, Wanying Ge +5

Music Source Restoration (MSR) extends source separation to realistic settings where signals undergo production effects (equalization, compression, reverb) and real-world degradati…