11 papers
Joint Text-Audio Alignment for EEG-to-Text Decoding in Chinese Speech Production and Perception
Tian Zheng, Xurong Xie, Xinxin Zhu +2
Decoding speech information directly from scalp electroencephalography (EEG) into text provides a potential non-invasive neural communication pathway for individuals with severe sp…
Can MLLMs Reason Beyond Language? VisReason: A Comprehensive Benchmark for Vision-Centric Reasoning
Longteng Guo, Yifan Wang, Pengkang Huo +4
Recent multimodal large language models (MLLMs) achieve strong performance on visual reasoning benchmarks, yet it remains unclear to what extent such performance reflects reasoning…
AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding
Handong Li, Zikang Liu, Longteng Guo +10
Processing long-form videos with Video Large Language Models (Video-LLMs) is computationally prohibitive. Current efficiency methods often compromise fine-grained perception throug…
SAGE: Multi-Agent Self-Evolution for LLM Reasoning
Yulin Peng, Xinxin Zhu, Chenxing Wei +4
Reinforcement learning with verifiable rewards improves reasoning in large language models (LLMs), but many methods still rely on large human-labeled datasets. While self-play redu…
SEMAG: Self-Evolutionary Multi-Agent Code Generation
Yulin Peng, Haowen Hou, Xinxin Zhu +2
Large Language Models (LLMs) have made significant progress in handling complex programming tasks. However, current methods rely on manual model selection and fixed workflows, whic…
ProAV-DiT: A Projected Latent Diffusion Transformer for Efficient Synchronized Audio-Video Generation
Jiahui Sun, Weining Wang, Mingzhen Sun +3
Sounding Video Generation (SVG) remains a challenging task due to the inherent structural misalignment between audio and video, as well as the high computational cost of multimodal…