computer-use benchmarking 1interactive agents 1long-horizon tasks 1safety auditing 1tool-use evaluation 1
From the 1 of 7 linked papers with an AI index.
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs
Ziqiao Shang, Lingyue Ge, Ling-Yue Ge +12
Systematically evaluating Multimodal Large Language Models (MLLMs) is essential for advancing Artificial General Intelligence (AGI). Yet existing benchmarks remain inadequate for r…
cs.LG2025
Generalizing Teacher Networks for Effective Knowledge Distillation Across Student Architectures
Kuluhan Binici, Weiming Wu, Tulika Mitra
Knowledge distillation (KD) is a model compression method that entails training a compact student model to emulate the performance of a more complex teacher model. However, the arc…