5 papers · 1 filter
Next-Token Prediction Task Assumes Optimal Data Ordering for LLM Training in Proof Generation
Chenyang An, Shima Imani, Feng Yao +8
In the field of large language model (LLM)-based proof generation, despite extensive training on large datasets such as ArXiv, LLMs still exhibit only modest performance on proving…
Linear Correlation in LM's Compositional Generalization and Hallucination
Letian Peng, Chenyang An, Shibo Hao +2
The generalization of language models (LMs) is undergoing active debates, contrasting their potential for general intelligence with their struggles with basic knowledge composition…
When is the consistent prediction likely to be a correct prediction?
Alex Nguyen, Dheeraj Mekala, Chengyu Dong +1
Self-consistency (Wang et al., 2023) suggests that the most consistent answer obtained through large language models (LLMs) is more likely to be correct. In this paper, we challeng…
Text Grafting: Near-Distribution Weak Supervision for Minority Classes in Text Classification
Letian Peng, Yi Gu, Chengyu Dong +2
For extremely weak-supervised text classification, pioneer research generates pseudo labels by mining texts similar to the class names from the raw corpus, which may end up with ve…
Evaluating the Smooth Control of Attribute Intensity in Text Generation with LLMs
Shang Zhou, Feng Yao, Chengyu Dong +2
Controlling the attribute intensity of text generation is crucial across scenarios (e.g., writing conciseness, chatting emotion, and explanation clarity). The remarkable capabiliti…