From the 1 of 4 linked papers with an AI index.
4 papers
NeMo: Needle in a Montage for Video-Language Understanding
Zi-Yuan Hu, Shuo Liang, Duo Zheng +10
The paper introduces the Needle in a Montage (NeMo) task and the NeMoBench benchmark to evaluate temporal understanding in video-language models, using an automated pipeline to gen…
Optimal Bayesian Stopping for Efficient Inference of Consistent LLM Answers
Jingkai Huang, Will Ma, Zhengyuan Zhou
A simple strategy for improving LLM accuracy, especially in math and reasoning problems, is to sample multiple responses and submit the answer most consistently reached. In this pa…
BEAR: Budgeted Evidence Allocation for Multi-Document Reasoning
Lin Sun, Linglin Zhang, Jingang Huang +3
We argue that multi-document reasoning is constrained not only by how much text a model can read, but also by how limited query-time evidence budget is allocated across documents a…
MagicWand: A Universal Agent for Generation and Evaluation Aligned with User Preference
Zitong Xu, Dake Shen, Yaosong Du +3
Recent advances in AIGC (Artificial Intelligence Generated Content) models have enabled significant progress in image and video generation. However, users still struggle to obtain…