collaborators

5 papers

cs.SD2026

Closing the Verification Loop: Self-Check Captioning for Long-Paragraph Detailed Audio Captioning

Fengji Ma, Yan Rong, Xu Li +3

Long-paragraph detailed audio captioning, which requires dense and transcript-faithful descriptions of fine-grained audio content, remains unsolved for current audio-visual multimo…

cs.SD2026

ACE-Cap: Active Evidence Acquisition via Agentic Co-Evolution for Long-Paragraph Fine-Grained Audio Captioning

Fengji Ma, Yan Rong, Xu Li +3

Long-paragraph fine-grained audio captioning requires models to recover diverse acoustic facts while avoiding omissions and unsupported details. However, prevailing captioners rema…

cs.SD2026

AudioMap: Cloze-and-Choice Reinforcement Learning for Time-Aware Dense Audio Captioning

Yan Rong, Fengji Ma, Xu Li +3

Time-aware dense audio captioning (TDAC) aims to generate multiple fine-grained attributes (dense) of the audio with precise time boundaries (time-aware). Existing methods struggle…

cs.SD2026

AudioScape-TTA: A Structured Soundscape Benchmark for Fine-Grained Text-to-Audio Evaluation

Jinting Wang, Yuguang Yang, Shengyu Li +4

Text-to-audio (TTA) generation has recently achieved remarkable progress in synthesizing realistic audio from natural language descriptions. However, determining whether generated…

cs.SD2025

PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation

Tianxin Xie, Wentao Lei, Kai Jiang +27

Text-to-audio-video (T2AV) generation is central to applications such as filmmaking and world modeling. However, current models often fail to produce physically plausible sounds. P…