Publications (14)
Mitigating Judgment Preference Bias in Large Language Models through Group-Based Polling
Shuliang Liu, Zhipeng Xu, Zhenghao Liu +6
Large Language Models (LLMs) as automatic evaluators, commonly referred to as LLM-as-a-Judge, have also attracted growing attention. This approach plays a vital role in aligning LL…
AIBench: An Agile Domain-specific Benchmarking Methodology and an AI Benchmark Suite
Wanling Gao, Fei Tang, Jianfeng Zhan +31
Domain-specific software and hardware co-design is encouraging as it is much easier to achieve efficiency for fewer tasks. Agile domain-specific benchmarking speeds up the process…
Advancing Knowledge Tracing by Exploring Follow-up Performance Trends
Hengyu Liu, Yushuai Li, Minghe Yu +6
Intelligent Tutoring Systems (ITS), such as Massive Open Online Courses, offer new opportunities for human learning. At the core of such systems, knowledge tracing (KT) predicts st…
MEGRAG: Multi-Granular Evidence Graphs for Answer-Aware Multi-Hop RAG
Weidong Bao, Yingying Sun, Jun Yang +7
Multi-hop question answering is a fundamental challenge in retrieval-augmented generation (RAG), because deriving an answer requires integrating dispersed evidence. Iterative RAG (…
COAST: Enhancing the Code Debugging Ability of LLMs through Communicative Agent Based Data Synthesis
Weiqing Yang, Hanbin Wang, Zhenghao Liu +7
Code debugging is a vital stage of software development, essential for ensuring the reliability and performance of Large Language Models (LLMs) in the code generation task. Human d…
AIBench: An Industry Standard Internet Service AI Benchmark Suite
Wanling Gao, Fei Tang, Lei Wang +22
Today's Internet Services are undergoing fundamental changes and shifting to an intelligent computing era where AI is widely employed to augment services. In this context, many inn…
Mixture-of-Retrieval Experts for Reasoning-Guided Multimodal Knowledge Exploitation
Chunyi Peng, Zhipeng Xu, Zhenghao Liu +7
Multimodal Retrieval-Augmented Generation (MRAG) has shown promise in mitigating hallucinations in Multimodal Large Language Models (MLLMs) by incorporating external knowledge. How…
Automated Formalization via Conceptual Retrieval-Augmented LLMs
Wangyue Lu, Lun Du, Sirui Li +6
Interactive theorem provers (ITPs) require manual formalization, which is labor-intensive and demands expert knowledge. While automated formalization offers a potential solution, i…
Interpretable Knowledge Tracing via Response Influence-based Counterfactual Reasoning
Jiajun Cui, Minghe Yu, Bo Jiang +3
Knowledge tracing (KT) plays a crucial role in computer-aided education and intelligent tutoring systems, aiming to assess students' knowledge proficiency by predicting their futur…
AIBench Training: Balanced Industry-Standard AI Training Benchmarking
Fei Tang, Wanling Gao, Jianfeng Zhan +30
Earlier-stage evaluations of a new AI architecture/system need affordable benchmarks. Only using a few AI component benchmarks like MLPerfalone in the other stages may lead to misl…
Enhancing Long-Chain Reasoning Distillation through Error-Aware Self-Reflection
Zhuoyang Wu, Xinze Li, Zhenghao Liu +7
Large Language Models (LLMs) have exhibited strong reasoning capabilities and achieved remarkable performance in mathematical problem-solving tasks. Recently, distilling reasoning…
A Probabilistic Generative Model for Tracking Multi-Knowledge Concept Mastery Probability
Hengyu Liu, Tiancheng Zhang, Fan Li +2
Knowledge tracing aims to track students' knowledge status over time to predict students' future performance accurately. Markov chain-based knowledge tracking (MCKT) models can tra…
MSC-180: A Benchmark for Automated Formal Theorem Proving from Mathematical Subject Classification
Sirui Li, Wangyue Lu, Xiaorui Shi +7
Automated Theorem Proving (ATP) represents a core research direction in artificial intelligence for achieving formal reasoning and verification, playing a significant role in advan…
Finding What Matters: Anchoring Context Knowledge with Evolving Indices for Iterative Retrieval
Mingyan Wu, Zhenghao Liu, Xinze Li +7
Retrieval-Augmented Generation (RAG) has become a dominant paradigm for mitigating hallucinations in Large Language Models (LLMs) by incorporating external knowledge. However, exis…