collaborators

22 papers

cs.AI2026

Mendel Gödel Machine: Recursive Self-Improving Coding Agents via Comparative Evolution

Changzhi Liu, Yilun Liu, Sikuan Yan +2

Self-improving coding agents that iteratively rewrite their own source code have demonstrated impressive performance on coding tasks. However, existing solutions generally derive s…

cs.CL2026

Behavioral and Representational Evidence of Binomial Ordering Preferences in Large Language Models

Zhiqing Yang, Yilun Liu, Yunpu Ma +2

Large language models (LLMs) can readily reproduce conventional expressions, yet their ability to model gradient frequency distributions remains underexplored. We investigate this…

cs.CL2026

MADE: Beyond Scoring via a Multilingual Agentic Diagnosing Engine for Fine-Grained Evaluation Insights

Yilun Liu, Miao Zhang, Shimin Tao +9

Multilingual and multicultural benchmarks now cover dozens of languages and model families, but the resulting score landscapes remain metric-rich and insight-poor, necessitating fi…

cs.MA2026

PyraVid: Hierarchical Multimodal Memory for Long-Horizon Video Reasoning

Sikuan Yan, Sicheng Dong, Haotong Wang +8

Memory has become an increasingly important component of agentic systems, as these systems are expected to reason over long-term experience. However, prior work has largely focused…

cs.CL2026

M-DaQ: Retrieving Samples with Multilingual Diversity and Quality for Instruction Fine-Tuning Datasets

Chunguang Zhao, Yilun Liu, Pufan Zeng +10

Multilingual instruction fine-tuning (IFT) empowers large language models to generalize across diverse linguistic and cultural contexts; however, high-quality, systematically curat…

cs.CL2026

The GaoYao Benchmark: A Comprehensive Framework for Evaluating Multilingual and Multicultural Abilities of Large Language Models

Yilun Liu, Chunguang Zhao, Mengyao Piao +14

Evaluating the multilingual and multicultural capabilities of Large Language Models (LLMs) is essential for their global utility. However, current benchmarks face three critical li…