#benchmark evaluation
32 papers · 1 filter
Measurement Risk in Supervised Financial NLP: Rubric and Metric Sensitivity on JF-ICR
Sidi Chang, Peiying Zhu, Yuxiao Chen +1
The paper investigates how the wording of evaluation rubrics and the choice of metrics affect the reliability of supervised financial NLP benchmarks, using a Japanese implicit‑comm…
Not All Needles Are Found: How Fact Distribution and Don't Make It Up Prompts Shape Retrieval, Reasoning, and Hallucination in Long-Context LLMs
Amirali Ebrahimzadeh, Seyyed M. Salili
The paper studies how the placement of facts in large text corpora and the use of anti‑hallucination prompts affect the ability of long‑context language models to retrieve informat…
Image Matching Filtering and Refinement by Planes and Beyond
Fabio Bellavia, Zhenjun Zhao, Luca Morelli +1
The paper evaluates state‑of‑the‑art filtering and refinement techniques for image matching, introduces a new method that combines planar constraints with cross‑correlation, and sh…
Baselines Before Architecture: Evaluating Coding Agents for Autonomous Penetration Testing
Ananda Dhakal, Krish Neupane, Aarjan Chaudhary
The paper conducts a controlled evaluation of plain coding agents versus specialized penetration‑testing architectures on the XBOW benchmark, showing that baseline agents already s…
Stage-Level Executor Allocation in Apache Spark with Cost-Performance Trade-offs
Miriam Rateike, Isaac Waweru Wambugu, Celia Cintas +3
The paper introduces a stage‑level executor allocation method for serverless Apache Spark that uses tree‑ensemble models to predict runtime and cost, letting users specify cost‑per…
SCALECUA: Scaling Computer Use Agents with Verifiable Task Synthesis and Efficient Online RL
Bowen Lv, Xiao Liu, Yanyu Ren +7
The paper introduces ScaleCUA, a framework that generates verifiable tasks and improves online reinforcement learning efficiency for computer use agents, achieving state-of-the-art…