From the 1 of 14 linked papers with an AI index.
14 papers
MemDefrag: Latent Memory Defragmentation for Large Language Models
Ruiyi Yan, Zhuoyuan Mao, Yiwen Guo
The paper introduces MemDefrag, a training-free, model-agnostic method that uses middle-layer attention signals to reorder and filter latent memory fragments in large language mode…
SemanticAudio: Audio Generation and Editing in Semantic Space
Zheqi Dai, Guangyan Zhang, Haolin He +5
In recent years, Text-to-Audio Generation has achieved remarkable progress, offering sound creators powerful tools to transform textual inspirations into vivid audio. However, exis…
One-Step Token-to-Waveform Generation with MeanFlow in Latent Space
Zheqi Dai, Guangyan Zhang, Zhen Ye +5
Neural audio codecs are central to modern LLM-based Text-to-Speech (TTS) and multimodal systems. As low-bitrate semantic codecs gain prominence, the Token-to-Waveform (Token2Wav) d…
UniSRM: A Unified Speech Reward Model for Reasoning-Based Fine-grained Assessment
Yuanyuan Wang, Dongchao Yang, Yayue Deng +4
Evaluating speech generation still relies heavily on human judgments, such as Mean Opinion Score (MOS), which are expensive, subjective, and difficult to reproduce at scale. While…
DualOptim+: Bridging Shared and Decoupled Optimizer States for Better Machine Unlearning in Large Language Models
Xuyang Zhong, Qizhang Li, Yiwen Guo +1
We propose DualOptim+, a novel optimization framework for improving machine unlearning in large language models. It introduces a base state to capture common representations shared…
AudioRole: An Audio Dataset for Character Role-Playing in Large Language Models
Wenyu Li, Xiaoqi Jiao, Yi Chang +2
The creation of high-quality multimodal datasets remains fundamental for advancing role-playing capabilities in large language models (LLMs). While existing works predominantly foc…