5 papers
RankLLM: Weighted Ranking of LLMs by Quantifying Question Difficulty
Ziqian Zhang, Xingjian Hu, Yue Huang +8
Benchmarks establish a standardized evaluation framework to systematically assess the performance of large language models (LLMs), facilitating objective comparisons and driving ad…
SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded Tasks
Yilun Zhao, Kaiyan Zhang, Tiansheng Hu +15
We present SciArena, an open and collaborative platform for evaluating foundation models on scientific literature-grounded tasks. Unlike traditional benchmarks for scientific liter…
RTV-Bench: Benchmarking MLLM Continuous Perception, Understanding and Reasoning through Real-Time Video
Shuhang Xun, Sicheng Tao, Jungang Li +11
Multimodal Large Language Models (MLLMs) have made rapid progress in perception, understanding, and reasoning, yet existing benchmarks fall short in evaluating these abilities unde…
SPOGW: a Score-based Preference Optimization method via Group-Wise comparison for workflows
Yitong Cui, Liu Liu, Baosheng Yu +5
Large language models (LLMs) have exhibited significant capabilities in addressing challenging problems throughout various fields, often through the use of agentic workflows that a…
CLIP-Powered Domain Generalization and Domain Adaptation: A Comprehensive Survey
Jindong Li, Yongguang Li, Yali Fu +4
As machine learning evolves, domain generalization (DG) and domain adaptation (DA) have become crucial for enhancing model robustness across diverse environments. Contrastive Langu…