9 citations · 9 across the 25 of their papers we have counts for
37 papers
GenRubric: Self-Evolving Rubric Generation for Scalable LLM Evaluation
Yifan Chen, Haitao Li, Qingyao Ai +4
Large language models are increasingly used as scalable evaluators for open-ended tasks. However, many LLM judges derive query-specific criteria during scoring, leaving the evaluat…
Fetch-then-Explore: Decoupling Selection from Extraction over a Persistent Workspace for Search Agents
Qi Liu, Yiqun Chen, Zidan Chen +6
Search agents now answer questions that take dozens of searches to settle, yet how such an agent reads a page has drawn far less attention than how it finds one. Nearly all of them…
Diagnosing Search Behavior and Failure Modes in Long-Horizon Search Agents
Qi Liu, Jiaxin Mao, Fengbin Zhu +1
Deep search agents answer difficult information-seeking questions by iteratively issuing search queries to gather supporting evidence, but it remains unclear whether and how greate…
FinDeepIndicator: Benchmarking Deep Research Agents in End-to-End Financial Indicator Construction
Chaoqun Yang, Fengbin Zhu, Xinyu Lin +5
Financial indicators are essential tools for transforming raw financial data into interpretable measures for various downstream tasks, such as valuation, risk assessment, and econo…
Herculean: An Agentic Benchmark for Financial Intelligence
Xueqing Peng, Zhuohan Xie, Yupeng Cao +60
As AI agents improve, the central question is no longer whether they can solve isolated well-defined financial tasks, but whether they can reliably carry out financial professional…
Plant, Persist, Trigger: Sleeper Attack on Large Language Model Agents
Yongxiang Li, Moxin Li, Zhixin Ma +4
Large Language Model (LLM) agents remain vulnerable to safety threats from the external environment, where attackers inject adversarial content into external observations such as t…