4 papers
Many-Shot CoT-ICL: Making In-Context Learning Truly Learn
Tsz Ting Chung, Lemao Liu, Mo Yu +1
While many-shot ICL achieves remarkable performance, prior studies of its scaling behavior have mainly focused on non-reasoning tasks. In this work, we study many-shot ICL on reaso…
Judge Like Human Examiners: A Weighted Importance Multi-Point Evaluation Framework for Generative Tasks with Long-form Answers
Guoxin Yu, Chulun Zhou, Lemao Liu +7
Evaluating the quality of model responses remains challenging in generative tasks with long-form answers, as the expected answers usually contain multiple semantically distinct yet…
DivLogicEval: A Framework for Benchmarking Logical Reasoning Evaluation in Large Language Models
Tsz Ting Chung, Lemao Liu, Mo Yu +1
Logic reasoning in natural language has been recognized as an important measure of human intelligence for Large Language Models (LLMs). Popular benchmarks may entangle multiple rea…
Large Language Models Can Self-Improve in Long-context Reasoning
Siheng Li, Cheng Yang, Zesen Cheng +4
Large language models (LLMs) have achieved substantial progress in processing long contexts but still struggle with long-context reasoning. Existing approaches typically involve fi…