collaborators

5 papers

cs.CL2026

ELBench: A Multi-Dimensional Benchmark for Education-Facing Large Language Models

Yilin Jiang, Xiaorong Zhu, Fei Tan +9

Large language models are increasingly deployed in education as tutors, teaching assistants, and content generators. These roles place demands that ordinary question answering does…

econ.GN2026

On Benchmark Hacking in ML Contests: Modeling, Insights and Design

Xiaoyun Qiu, Yang Yu, Haifeng Xu

Benchmark hacking refers to tuning a machine learning model to score highly on certain evaluation criteria without improving true generalization or faithfully solving the intended…

cs.AI2026

InfiCoEvalChain: A Blockchain-Based Decentralized Framework for Collaborative LLM Evaluation

Yifan Yang, Jinjia Li, Kunxi Li +7

The rapid advancement of large language models (LLMs) demands increasingly reliable evaluation, yet current centralized evaluation suffers from opacity, overfitting, and hardware-i…

cs.AI2025

InfiAgent: Self-Evolving Pyramid Agent Framework for Infinite Scenarios

Chenglin Yu, Yang Yu, Songmiao Wang +5

Large Language Model (LLM) agents have demonstrated remarkable capabilities in organizing and executing complex tasks, and many such agents are now widely used in various applicati…

cs.CL2025

InfiR : Crafting Effective Small Language Models and Multimodal Small Language Models in Reasoning

Congkai Xie, Shuo Cai, Wenjun Wang +17

Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs) have made significant advancements in reasoning capabilities. However, they still face challenges such as…