4 citations · 12 across the 40 of their papers we have counts for
13 papers · 1 filter
How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks
Ali Ansari, Haoran Sun, Andy Zeyi Liu +48
Low reported scores on leading physics benchmarks, including those featured in the Artificial Analysis Intelligence Index (2026), suggest that frontier language models still strugg…
Herculean: An Agentic Benchmark for Financial Intelligence
Xueqing Peng, Zhuohan Xie, Yupeng Cao +60
As AI agents improve, the central question is no longer whether they can solve isolated well-defined financial tasks, but whether they can reliably carry out financial professional…
Step-level Optimization for Efficient Computer-use Agents
Jinbiao Wei, Kangqi Ni, Yilun Zhao +2
Computer-use agents provide a promising path toward general software automation because they can interact directly with arbitrary graphical user interfaces instead of relying on br…
A Survey of Multimodal Mathematical Reasoning: From Perception, Alignment to Reasoning
Tianyu Yang, Sihong Wu, Yilun Zhao +6
Multimodal Mathematical Reasoning (MMR) has recently attracted increasing attention for its capability to solve mathematical problems involving both textual and visual modalities.…
ANCHOR: Branch-Point Data Generation for GUI Agents
Jinbiao Wei, Yilun Zhao, Kangqi Ni +1
End-to-end GUI agents for real desktop environments require large amounts of high-quality interaction data, yet collecting human demonstrations is expensive and existing synthetic…
PuzzlePlex: Benchmarking Foundation Models on Reasoning and Planning with Puzzles
Yitao Long, Yuru Jiang, Hongjun Liu +6
This work investigates the reasoning and planning capabilities of foundation models and their scalability in complex, dynamic environments. We introduce PuzzlePlex, a benchmark des…