6 papers
Benchmarking at the Edge of Comprehension
Samuele Marro, Jialin Yu, Emanuele La Malfa +8
As frontier Large Language Models (LLMs) increasingly saturate new benchmarks shortly after they are published, benchmarking itself is at a juncture: if frontier models keep improv…
HoloBrain-0 Technical Report
Xuewu Lin, Tianwei Lin, Yun Du +12
In this work, we introduce HoloBrain-0, a comprehensive Vision-Language-Action (VLA) framework that bridges the gap between foundation model research and reliable real-world robot…
Collaborative Belief Reasoning with LLMs for Efficient Multi-Agent Collaboration
Zhimin Wang, Duo Wu, Shaokang He +6
Effective real-world multi-agent collaboration requires not only accurate planning but also the ability to reason about collaborators' intents--a crucial capability for avoiding mi…
Detecting Instruction Fine-tuning Attacks using Influence Function
Jiawei Li
Instruction fine-tuning attacks pose a serious threat to large language models (LLMs) by subtly embedding poisoned examples in fine-tuning datasets, leading to harmful or unintende…
Delta-Influence: Unlearning Poisons via Influence Functions
Wenjie Li, Jiawei Li, Pengcheng Zeng +3
Addressing data integrity challenges, such as unlearning the effects of data poisoning after model training, is necessary for the reliable deployment of machine learning models. St…
Predicting User Behavior in Smart Spaces with LLM-Enhanced Logs and Personalized Prompts
Yunpeng Song, Jiawei Li, Yiheng Bian +1
Enhancing the intelligence of smart systems, such as smart home, and smart vehicle, and smart grids, critically depends on developing sophisticated planning capabilities that can a…