#benchmark evaluation

topicbenchmark evaluation

32 papers · 1 filter

cs.AI2026

Measurement Risk in Supervised Financial NLP: Rubric and Metric Sensitivity on JF-ICR

Sidi Chang, Peiying Zhu, Yuxiao Chen +1

The paper investigates how the wording of evaluation rubrics and the choice of metrics affect the reliability of supervised financial NLP benchmarks, using a Japanese implicit‑comm…

cs.CL2026

Not All Needles Are Found: How Fact Distribution and Don't Make It Up Prompts Shape Retrieval, Reasoning, and Hallucination in Long-Context LLMs

Amirali Ebrahimzadeh, Seyyed M. Salili

The paper studies how the placement of facts in large text corpora and the use of anti‑hallucination prompts affect the ability of long‑context language models to retrieve informat…

cs.CV2026

Image Matching Filtering and Refinement by Planes and Beyond

Fabio Bellavia, Zhenjun Zhao, Luca Morelli +1

The paper evaluates state‑of‑the‑art filtering and refinement techniques for image matching, introduces a new method that combines planar constraints with cross‑correlation, and sh…

cs.CR2026

Baselines Before Architecture: Evaluating Coding Agents for Autonomous Penetration Testing

Ananda Dhakal, Krish Neupane, Aarjan Chaudhary

The paper conducts a controlled evaluation of plain coding agents versus specialized penetration‑testing architectures on the XBOW benchmark, showing that baseline agents already s…

cs.DC2026

Stage-Level Executor Allocation in Apache Spark with Cost-Performance Trade-offs

Miriam Rateike, Isaac Waweru Wambugu, Celia Cintas +3

The paper introduces a stage‑level executor allocation method for serverless Apache Spark that uses tree‑ensemble models to predict runtime and cost, letting users specify cost‑per…

cs.AI2026

SCALECUA: Scaling Computer Use Agents with Verifiable Task Synthesis and Efficient Online RL

Bowen Lv, Xiao Liu, Yanyu Ren +7

The paper introduces ScaleCUA, a framework that generates verifiable tasks and improves online reinforcement learning efficiency for computer use agents, achieving state-of-the-art…