papers

Publications (14)

cs.CL2026

Mitigating Judgment Preference Bias in Large Language Models through Group-Based Polling

Shuliang Liu, Zhipeng Xu, Zhenghao Liu +6

Large Language Models (LLMs) as automatic evaluators, commonly referred to as LLM-as-a-Judge, have also attracted growing attention. This approach plays a vital role in aligning LL…

cs.PF2020

AIBench: An Agile Domain-specific Benchmarking Methodology and an AI Benchmark Suite

Wanling Gao, Fei Tang, Jianfeng Zhan +31

Domain-specific software and hardware co-design is encouraging as it is much easier to achieve efficiency for fewer tasks. Agile domain-specific benchmarking speeds up the process…

cs.CY2025

Advancing Knowledge Tracing by Exploring Follow-up Performance Trends

Hengyu Liu, Yushuai Li, Minghe Yu +6

Intelligent Tutoring Systems (ITS), such as Massive Open Online Courses, offer new opportunities for human learning. At the core of such systems, knowledge tracing (KT) predicts st…

cs.AI2026

MEGRAG: Multi-Granular Evidence Graphs for Answer-Aware Multi-Hop RAG

Weidong Bao, Yingying Sun, Jun Yang +7

Multi-hop question answering is a fundamental challenge in retrieval-augmented generation (RAG), because deriving an answer requires integrating dispersed evidence. Iterative RAG (…

cs.SE2025

COAST: Enhancing the Code Debugging Ability of LLMs through Communicative Agent Based Data Synthesis

Weiqing Yang, Hanbin Wang, Zhenghao Liu +7

Code debugging is a vital stage of software development, essential for ensuring the reliability and performance of Large Language Models (LLMs) in the code generation task. Human d…

cs.CV2019

AIBench: An Industry Standard Internet Service AI Benchmark Suite

Wanling Gao, Fei Tang, Lei Wang +22

Today's Internet Services are undergoing fundamental changes and shifting to an intelligent computing era where AI is widely employed to augment services. In this context, many inn…

cs.CL2026

Mixture-of-Retrieval Experts for Reasoning-Guided Multimodal Knowledge Exploitation

Chunyi Peng, Zhipeng Xu, Zhenghao Liu +7

Multimodal Retrieval-Augmented Generation (MRAG) has shown promise in mitigating hallucinations in Multimodal Large Language Models (MLLMs) by incorporating external knowledge. How…

cs.AI2026

Automated Formalization via Conceptual Retrieval-Augmented LLMs

Wangyue Lu, Lun Du, Sirui Li +6

Interactive theorem provers (ITPs) require manual formalization, which is labor-intensive and demands expert knowledge. While automated formalization offers a potential solution, i…

cs.CY2024

Interpretable Knowledge Tracing via Response Influence-based Counterfactual Reasoning

Jiajun Cui, Minghe Yu, Bo Jiang +3

Knowledge tracing (KT) plays a crucial role in computer-aided education and intelligent tutoring systems, aiming to assess students' knowledge proficiency by predicting their futur…

cs.AI2021

AIBench Training: Balanced Industry-Standard AI Training Benchmarking

Fei Tang, Wanling Gao, Jianfeng Zhan +30

Earlier-stage evaluations of a new AI architecture/system need affordable benchmarks. Only using a few AI component benchmarks like MLPerfalone in the other stages may lead to misl…

cs.CL2025

Enhancing Long-Chain Reasoning Distillation through Error-Aware Self-Reflection

Zhuoyang Wu, Xinze Li, Zhenghao Liu +7

Large Language Models (LLMs) have exhibited strong reasoning capabilities and achieved remarkable performance in mathematical problem-solving tasks. Recently, distilling reasoning…

cs.LG2023

A Probabilistic Generative Model for Tracking Multi-Knowledge Concept Mastery Probability

Hengyu Liu, Tiancheng Zhang, Fan Li +2

Knowledge tracing aims to track students' knowledge status over time to predict students' future performance accurately. Markov chain-based knowledge tracking (MCKT) models can tra…

cs.AI2025

MSC-180: A Benchmark for Automated Formal Theorem Proving from Mathematical Subject Classification

Sirui Li, Wangyue Lu, Xiaorui Shi +7

Automated Theorem Proving (ATP) represents a core research direction in artificial intelligence for achieving formal reasoning and verification, playing a significant role in advan…

cs.CL2026

Finding What Matters: Anchoring Context Knowledge with Evolving Indices for Iterative Retrieval

Mingyan Wu, Zhenghao Liu, Xinze Li +7

Retrieval-Augmented Generation (RAG) has become a dominant paradigm for mitigating hallucinations in Large Language Models (LLMs) by incorporating external knowledge. However, exis…