From the 1 of 21 linked papers with an AI index.
1 citations · 1 across the 4 of their papers we have counts for
13 papers · 1 filter
ANAH-v2: Scaling Analytical Hallucination Annotation of Large Language Models
Yuzhe Gu, Ziwei Ji, Wenwei Zhang +3
Large language models (LLMs) exhibit hallucinations in long-form question-answering tasks across various domains and wide applications. Current hallucination detection and mitigati…
Training Language Models to Critique With Multi-agent Feedback
Tian Lan, Wenwei Zhang, Chengqi Lyu +6
Critique ability, a meta-cognitive capability of humans, presents significant challenges for LLMs to improve. Recent works primarily rely on supervised fine-tuning (SFT) using crit…
CriticEval: Evaluating Large Language Model as Critic
Tian Lan, Wenwei Zhang, Chen Xu +4
Critique ability, i.e., the capability of Large Language Models (LLMs) to identify and rectify flaws in responses, is crucial for their applications in self-improvement and scalabl…
Scaling Behavior for Large Language Models regarding Numeral Systems: An Example using Pythia
Zhejian Zhou, Jiayu Wang, Dahua Lin +1
Though Large Language Models (LLMs) have shown remarkable abilities in mathematics reasoning, they are still struggling with performing numeric operations accurately, such as addit…
LEAN-GitHub: Compiling GitHub LEAN repositories for a versatile LEAN prover
Zijian Wu, Jiayu Wang, Dahua Lin +1
Recently, large language models have presented promising results in aiding formal mathematical reasoning. However, their performance is restricted due to the scarcity of formal the…
InternLM-Law: An Open Source Chinese Legal Large Language Model
Zhiwei Fei, Songyang Zhang, Xiaoyu Shen +9
While large language models (LLMs) have showcased impressive capabilities, they struggle with addressing legal queries due to the intricate complexities and specialized expertise r…