activity
20242026
most citedA Survey on Speech Large Language Models for Understanding

6 citations · 8 across the 17 of their papers we have counts for

collaborators
Showing 2026 · cs.SDShow all

7 papers · 2 filters

cs.SD2026

Closing the Verification Loop: Self-Check Captioning for Long-Paragraph Detailed Audio Captioning

Fengji Ma, Yan Rong, Xu Li +3

Long-paragraph detailed audio captioning, which requires dense and transcript-faithful descriptions of fine-grained audio content, remains unsolved for current audio-visual multimo…

cs.SD2026

ACE-Cap: Active Evidence Acquisition via Agentic Co-Evolution for Long-Paragraph Fine-Grained Audio Captioning

Fengji Ma, Yan Rong, Xu Li +3

Long-paragraph fine-grained audio captioning requires models to recover diverse acoustic facts while avoiding omissions and unsupported details. However, prevailing captioners rema…

cs.SD2026

SingDance: Compositional Zero-Shot Singing-and-Dancing Video Generation with Role-Aware Audio Conditioning

Tao Feng, Xu Li, Xiangyang Luo +4

Generating personalized dance videos from a reference image, text prompt, and audio track requires music-conditioned body motion. Singing-and-dancing adds a second requirement: the…

cs.SD2026

SonicWeave: Chunk-Routed Mixture-of-Experts for Unified Audio Scene Generation

Yunrui Cai, Xu Li, Yucheng Zhou +8

Text-conditioned general audio generation is moving beyond isolated speech, music, and sound-effect synthesis toward a single model that can compose them into controllable, coheren…

cs.SD2026

AudioMap: Cloze-and-Choice Reinforcement Learning for Time-Aware Dense Audio Captioning

Yan Rong, Fengji Ma, Xu Li +3

Time-aware dense audio captioning (TDAC) aims to generate multiple fine-grained attributes (dense) of the audio with precise time boundaries (time-aware). Existing methods struggle…

cs.SD2026

FlexiSLM: A Spoken Language Model with Dynamic and Controllable Frame Rates

Jiaqi Li, Chaoren Wang, Xiaohai Tian +9

Spoken language models (SLMs) extend LLMs to speech input and output, but existing systems use fixed frame rates (e.g., 25 or 12.5 Hz), overlooking speech's time-varying informatio…