10 papers
CoLT: Reasoning with Chain of Latent Tool Calls
Fangwei Zhu, Zhifang Sui
Chain-of-Thought (CoT) is a critical technique in enhancing the reasoning ability of Large Language Models (LLMs), and latent reasoning methods have been proposed to accelerate the…
RICo: Refined In-Context Contribution for Automatic Instruction-Tuning Data Selection
Yixin Yang, Qingxiu Dong, Linli Yao +2
Data selection for instruction tuning is crucial for improving the performance of large language models (LLMs) while reducing training costs. In this paper, we propose Refined Cont…
Chain-of-Thought Tokens are Computer Program Variables
Fangwei Zhu, Peiyi Wang, Zhifang Sui
Chain-of-thoughts (CoT) requires large language models (LLMs) to generate intermediate steps before reaching the final answer, and has been proven effective to help LLMs solve comp…
Chip-Tuning: Classify Before Language Models Say
Fangwei Zhu, Dian Li, Jiajun Huang +3
The rapid development in the performance of large language models (LLMs) is accompanied by the escalation of model size, leading to the increasing cost of model training and infere…
LLMAEL: Large Language Models are Good Context Augmenters for Entity Linking
Amy Xin, Yunjia Qi, Zijun Yao +5
Specialized entity linking (EL) models are well-trained at mapping mentions to unique knowledge base (KB) entities according to a given context. However, specialized EL models stru…
CoUDA: Coherence Evaluation via Unified Data Augmentation
Dawei Zhu, Wenhao Wu, Yifan Song +3
Coherence evaluation aims to assess the organization and structure of a discourse, which remains challenging even in the era of large language models. Due to the scarcity of annota…