2 papers
cs.CL2025
Can LLMs Correct Themselves? A Benchmark of Self-Correction in LLMs
Guiyao Tie, Zenghui Yuan, Zeli Zhao +11
Self-correction of large language models (LLMs) emerges as a critical component for enhancing their reasoning performance. Although various self-correction methods have been propos…
cs.AI2025
MMLU-Reason: Benchmarking Multi-Task Multi-modal Language Understanding and Reasoning
Guiyao Tie, Xueyang Zhou, Tianhe Gu +7
Recent advances in Multi-Modal Large Language Models (MLLMs) have enabled unified processing of language, vision, and structured inputs, opening the door to complex tasks such as l…