11 papers
When LLM Judges Inflate Scores: Exploring Overrating in Relevance Assessment
Chuting Yu, Hang Li, Guido Zuccon +2
Human relevance assessment is time-consuming and cognitively intensive, limiting the scalability of Information Retrieval evaluation. This has led to growing interest in using larg…
Reasoned Safety Alignment: Ensuring Jailbreak Defense via Answer-Then-Check
Chentao Cao, Xiaojun Xu, Bo Han +1
As large language models (LLMs) continue to advance in capabilities, ensuring their safety against jailbreak attacks remains a critical challenge. In this paper, we introduce a nov…
How Uncertain Is the Grade? A Benchmark of Uncertainty Metrics for LLM-Based Automatic Assessment
Hang Li, Kaiqi Yang, Xianxuan Long +9
The rapid rise of large language models (LLMs) is reshaping the landscape of automatic assessment in education. While these systems demonstrate substantial advantages in adaptabili…
InternVLA-A1: Unifying Understanding, Generation and Action for Robotic Manipulation
Junhao Cai, Zetao Cai, Jiafei Cao +39
Prevalent Vision-Language-Action (VLA) models are typically built upon Multimodal Large Language Models (MLLMs) and demonstrate exceptional proficiency in semantic understanding, b…
Memory Retrieval and Consolidation in Large Language Models through Function Tokens
Shaohua Zhang, Yuan Lin, Hang Li
The remarkable success of large language models (LLMs) stems from their ability to consolidate vast amounts of knowledge into the memory during pre-training and to retrieve it from…
How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs
Andrew Estornell, Jean-Francois Ton, Muhammad Faaiz Taufiq +1
Large Language Models (LLMs) have achieved strong performance on a wide range of complex reasoning tasks, yet further gains are often possible by leveraging the complementary stren…