3 papers
q-fin.TR2026
AlphaForgeBench: Benchmarking End-to-End Trading Strategy Design with Large Language Models
Wentao Zhang, Mingxuan Zhao, Jincheng Gao +5
The rapid advancement of Large Language Models (LLMs) has led to a surge of financial benchmarks, evolving from static knowledge evaluation toward interactive trading simulations.…
cs.CV2026
Do Vision-Language Models Measure Up? Benchmarking Visual Measurement Reading with MeasureBench
Fenfen Lin, Yesheng Liu, Haiyu Xu +7
Reading measurement instruments is effortless for humans and requires relatively little domain expertise, yet it remains surprisingly challenging for current vision-language models…
cs.CL2025
Beyond Multiple Choice: Verifiable OpenQA for Robust Vision-Language RFT
Yesheng Liu, Hao Li, Haiyu Xu +9
Multiple-choice question answering (MCQA) has been a popular format for evaluating and reinforcement fine-tuning (RFT) of modern multimodal language models. Its constrained output…