activity
20242026
most citedChemDFM-R: A Chemical Reasoning LLM Enhanced with Atomized Chemical Knowledge

2 citations · 2 across the 4 of their papers we have counts for

collaborators
Showing 2024Show all

6 papers · 1 filter

cs.AI2024

AdaEAGLE: Optimizing Speculative Decoding via Explicit Modeling of Adaptive Draft Structures

Situo Zhang, Hankun Wang, Da Ma +4

Speculative Decoding (SD) is a popular lossless technique for accelerating the inference of Large Language Models (LLMs). We show that the decoding speed of SD frameworks with stat…

cs.CL2024

SciDFM: A Large Language Model with Mixture-of-Experts for Science

Liangtai Sun, Danyu Luo, Da Ma +7

Recently, there has been a significant upsurge of interest in leveraging large language models (LLMs) to assist scientific discovery. However, most LLMs only focus on general scien…

cs.CL2024

SciEval: A Multi-Level Large Language Model Evaluation Benchmark for Scientific Research

Liangtai Sun, Yang Han, Zihan Zhao +5

Recently, there has been growing interest in using Large Language Models (LLMs) for scientific research. Numerous benchmarks have been proposed to evaluate the ability of LLMs for…

cs.CL2024

Rejection Improves Reliability: Training LLMs to Refuse Unknown Questions Using RL from Knowledge Feedback

Hongshen Xu, Zichen Zhu, Situo Zhang +4

Large Language Models (LLMs) often generate erroneous outputs, known as hallucinations, due to their limitations in discerning questions beyond their knowledge scope. While address…

cs.CL2024

Evolving Subnetwork Training for Large Language Models

Hanqi Li, Lu Chen, Da Ma +3

Large language models have ushered in a new era of artificial intelligence research. However, their substantial training costs hinder further development and widespread adoption. I…

cs.CL2024

Sparsity-Accelerated Training for Large Language Models

Da Ma, Lu Chen, Pengyu Wang +6

Large language models (LLMs) have demonstrated proficiency across various natural language processing (NLP) tasks but often require additional training, such as continual pre-train…