4 papers
DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models
Luolin Xiong, Haofen Wang, Xi Chen +7
DeepSeek, a Chinese Artificial Intelligence (AI) startup, has released their V3 and R1 series models, which attracted global attention due to their low cost, high performance, and…
Hadamard Adapter: An Extreme Parameter-Efficient Adapter Tuning Method for Pre-trained Language Models
Yuyan Chen, Qiang Fu, Ge Fan +6
Recent years, Pre-trained Language models (PLMs) have swept into various fields of artificial intelligence and achieved great success. However, most PLMs, such as T5 and GPT3, have…
Hallucination Detection: Robustly Discerning Reliable Answers in Large Language Models
Yuyan Chen, Qiang Fu, Yichen Yuan +6
Large Language Models (LLMs) have gained widespread adoption in various natural language processing tasks, including question answering and dialogue systems. However, a major drawb…
Source Prompt: Coordinated Pre-training of Language Models on Diverse Corpora from Multiple Sources
Yipei Xu, Dakuan Lu, Jiaqing Liang +7
Pre-trained language models (PLMs) have established the new paradigm in the field of NLP. For more powerful PLMs, one of the most popular and successful way is to continuously scal…