collaborators

10 papers

cs.CL2026

Beyond Input Understanding: Diagnosing Multilingual Mathematical Reasoning with Directed Acyclic Trace Graphs

Jiaqiao Zhang, Zhoujun Li, Raoyuan Zhao +5

Large reasoning models (LRMs) achieve strong mathematical reasoning performance in English, but remain much less reliable in many low- and medium-resource languages. This gap is of…

cs.CL2026

ReverseMath: Answer Inversion for Scalable and Verifiable Mathematical Problem Generation

Raoyuan Zhao, Yihong Liu, Yupei Du +2

Mathematical reasoning benchmarks are vital for evaluating large language models (LLMs), but many are static and repeatedly exposed through public evaluation and training pipelines…

cs.CL2026

Crosslingual On-Policy Self-Distillation for Multilingual Reasoning

Yihong Liu, Raoyuan Zhao, Michael A. Hedderich +1

Large language models (LLMs) have achieved remarkable progress in mathematical reasoning, but this ability is not equally accessible across languages. Especially low-resource langu…

cs.CL2026

Evaluating Robustness of Large Language Models Against Multilingual Typographical Errors

Raoyuan Zhao, Yihong Liu, Lena Altinger +2

Large language models (LLMs) are increasingly deployed in multilingual, real-world applications with user inputs -- naturally introducing \emph{typographical errors} (typos). Yet m…

cs.CL2026

Large Reasoning Models Are (Not Yet) Multilingual Latent Reasoners

Yihong Liu, Raoyuan Zhao, Hinrich Schütze +1

Large reasoning models (LRMs) achieve strong performance on mathematical reasoning tasks, often attributed to their capability to generate explicit chain-of-thought (CoT) explanati…

cs.CV2025

Human Uncertainty-Aware Data Selection and Automatic Labeling in Visual Question Answering

Jian Lan, Zhicheng Liu, Udo Schlegel +5

Large vision-language models (VLMs) achieve strong performance in Visual Question Answering but still rely heavily on supervised fine-tuning (SFT) with massive labeled datasets, wh…