5 papers
ELBench: A Multi-Dimensional Benchmark for Education-Facing Large Language Models
Yilin Jiang, Xiaorong Zhu, Fei Tan +9
Large language models are increasingly deployed in education as tutors, teaching assistants, and content generators. These roles place demands that ordinary question answering does…
On Benchmark Hacking in ML Contests: Modeling, Insights and Design
Xiaoyun Qiu, Yang Yu, Haifeng Xu
Benchmark hacking refers to tuning a machine learning model to score highly on certain evaluation criteria without improving true generalization or faithfully solving the intended…
InfiCoEvalChain: A Blockchain-Based Decentralized Framework for Collaborative LLM Evaluation
Yifan Yang, Jinjia Li, Kunxi Li +7
The rapid advancement of large language models (LLMs) demands increasingly reliable evaluation, yet current centralized evaluation suffers from opacity, overfitting, and hardware-i…
InfiAgent: Self-Evolving Pyramid Agent Framework for Infinite Scenarios
Chenglin Yu, Yang Yu, Songmiao Wang +5
Large Language Model (LLM) agents have demonstrated remarkable capabilities in organizing and executing complex tasks, and many such agents are now widely used in various applicati…
InfiR : Crafting Effective Small Language Models and Multimodal Small Language Models in Reasoning
Congkai Xie, Shuo Cai, Wenjun Wang +17
Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs) have made significant advancements in reasoning capabilities. However, they still face challenges such as…