6 citations · 8 across the 17 of their papers we have counts for
7 papers · 2 filters
Closing the Verification Loop: Self-Check Captioning for Long-Paragraph Detailed Audio Captioning
Fengji Ma, Yan Rong, Xu Li +3
Long-paragraph detailed audio captioning, which requires dense and transcript-faithful descriptions of fine-grained audio content, remains unsolved for current audio-visual multimo…
ACE-Cap: Active Evidence Acquisition via Agentic Co-Evolution for Long-Paragraph Fine-Grained Audio Captioning
Fengji Ma, Yan Rong, Xu Li +3
Long-paragraph fine-grained audio captioning requires models to recover diverse acoustic facts while avoiding omissions and unsupported details. However, prevailing captioners rema…
SingDance: Compositional Zero-Shot Singing-and-Dancing Video Generation with Role-Aware Audio Conditioning
Tao Feng, Xu Li, Xiangyang Luo +4
Generating personalized dance videos from a reference image, text prompt, and audio track requires music-conditioned body motion. Singing-and-dancing adds a second requirement: the…
SonicWeave: Chunk-Routed Mixture-of-Experts for Unified Audio Scene Generation
Yunrui Cai, Xu Li, Yucheng Zhou +8
Text-conditioned general audio generation is moving beyond isolated speech, music, and sound-effect synthesis toward a single model that can compose them into controllable, coheren…
AudioMap: Cloze-and-Choice Reinforcement Learning for Time-Aware Dense Audio Captioning
Yan Rong, Fengji Ma, Xu Li +3
Time-aware dense audio captioning (TDAC) aims to generate multiple fine-grained attributes (dense) of the audio with precise time boundaries (time-aware). Existing methods struggle…
FlexiSLM: A Spoken Language Model with Dynamic and Controllable Frame Rates
Jiaqi Li, Chaoren Wang, Xiaohai Tian +9
Spoken language models (SLMs) extend LLMs to speech input and output, but existing systems use fixed frame rates (e.g., 25 or 12.5 Hz), overlooking speech's time-varying informatio…