22 papers
Mendel Gödel Machine: Recursive Self-Improving Coding Agents via Comparative Evolution
Changzhi Liu, Yilun Liu, Sikuan Yan +2
Self-improving coding agents that iteratively rewrite their own source code have demonstrated impressive performance on coding tasks. However, existing solutions generally derive s…
Behavioral and Representational Evidence of Binomial Ordering Preferences in Large Language Models
Zhiqing Yang, Yilun Liu, Yunpu Ma +2
Large language models (LLMs) can readily reproduce conventional expressions, yet their ability to model gradient frequency distributions remains underexplored. We investigate this…
MADE: Beyond Scoring via a Multilingual Agentic Diagnosing Engine for Fine-Grained Evaluation Insights
Yilun Liu, Miao Zhang, Shimin Tao +9
Multilingual and multicultural benchmarks now cover dozens of languages and model families, but the resulting score landscapes remain metric-rich and insight-poor, necessitating fi…
PyraVid: Hierarchical Multimodal Memory for Long-Horizon Video Reasoning
Sikuan Yan, Sicheng Dong, Haotong Wang +8
Memory has become an increasingly important component of agentic systems, as these systems are expected to reason over long-term experience. However, prior work has largely focused…
M-DaQ: Retrieving Samples with Multilingual Diversity and Quality for Instruction Fine-Tuning Datasets
Chunguang Zhao, Yilun Liu, Pufan Zeng +10
Multilingual instruction fine-tuning (IFT) empowers large language models to generalize across diverse linguistic and cultural contexts; however, high-quality, systematically curat…
The GaoYao Benchmark: A Comprehensive Framework for Evaluating Multilingual and Multicultural Abilities of Large Language Models
Yilun Liu, Chunguang Zhao, Mengyao Piao +14
Evaluating the multilingual and multicultural capabilities of Large Language Models (LLMs) is essential for their global utility. However, current benchmarks face three critical li…