most citedThe Price of Format: Diversity Collapse in LLMs

2 citations · 2 across the 4 of their papers we have counts for

collaborators

5 papers

cs.CL20252 cited

The Price of Format: Diversity Collapse in LLMs

Longfei Yun, Chenyang An, Zilong Wang +2

Instruction-tuned large language models (LLMs) employ structured templates, such as role markers and special tokens, to enforce format consistency during inference. However, we ide…

cs.CL2025

Linear Correlation in LM's Compositional Generalization and Hallucination

Letian Peng, Chenyang An, Shibo Hao +2

The generalization of language models (LMs) is undergoing active debates, contrasting their potential for general intelligence with their struggles with basic knowledge composition…

cs.CV2024

Diffusion-Augmented Coreset Expansion for Scalable Dataset Distillation

Ali Abbasi, Shima Imani, Chenyang An +6

With the rapid scaling of neural networks, data storage and communication demands have intensified. Dataset distillation has emerged as a promising solution, condensing information…

cs.CL2024

Next-Token Prediction Task Assumes Optimal Data Ordering for LLM Training in Proof Generation

Chenyang An, Shima Imani, Feng Yao +8

In the field of large language model (LLM)-based proof generation, despite extensive training on large datasets such as ArXiv, LLMs still exhibit only modest performance on proving…

cs.CL2024

Correlation and Navigation in the Vocabulary Key Representation Space of Language Models

Letian Peng, Chenyang An, Jingbo Shang

Language model (LM) decoding is based on the next-token prediction (NTP) probability distribution. For neural LMs (e.g., Transformer-based), NTP distribution is essentially a softm…