12 papers
Long Term Memory: The Foundation of AI Self-Evolution
Xun Jiang, Feng Li, Han Zhao +12
Large language models (LLMs) like GPTs, trained on vast datasets, have demonstrated impressive capabilities in language understanding, reasoning, and planning, achieving human-leve…
Delusions of Large Language Models
Hongshen Xu, Zixv yang, Zichen Zhu +7
Large Language Models often generate factually incorrect but plausible outputs, known as hallucinations. We identify a more insidious phenomenon, LLM delusion, defined as high beli…
SemantiCodec: An Ultra Low Bitrate Semantic Audio Codec for General Sound
Haohe Liu, Xuenan Xu, Yi Yuan +3
Large language models (LLMs) have significantly advanced audio processing through audio codecs that convert audio into discrete tokens, enabling the application of language modelli…
Unified Pathological Speech Analysis with Prompt Tuning
Fei Yang, Xuenan Xu, Mengyue Wu +1
Pathological speech analysis has been of interest in the detection of certain diseases like depression and Alzheimer's disease and attracts much interest from researchers. However,…
Auto-ACD: A Large-scale Dataset for Audio-Language Representation Learning
Luoyi Sun, Xuenan Xu, Mengyue Wu +1
Recently, the AI community has made significant strides in developing powerful foundation models, driven by large-scale multimodal datasets. However, for audio representation learn…
DiveSound: LLM-Assisted Automatic Taxonomy Construction for Diverse Audio Generation
Baihan Li, Zeyu Xie, Xuenan Xu +5
Audio generation has attracted significant attention. Despite remarkable enhancement in audio quality, existing models overlook diversity evaluation. This is partially due to the l…