1 citations · 1 across the 7 of their papers we have counts for
10 papers
SAGE: Benchmarking and Improving Retrieval for Deep Research Agents
Tiansheng Hu, Yilun Zhao, Canyu Zhang +2
Deep research agents have emerged as powerful systems for addressing complex queries. Meanwhile, LLM-based retrievers have demonstrated strong capability in following instructions…
LimRank: Less is More for Reasoning-Intensive Information Reranking
Tingyu Song, Yilun Zhao, Siyue Zhang +2
Existing approaches typically rely on large-scale fine-tuning to adapt LLMs for information reranking tasks, which is computationally expensive. In this work, we demonstrate that m…
FinTrust: A Comprehensive Benchmark of Trustworthiness Evaluation in Finance Domain
Tiansheng Hu, Tongyan Hu, Liuyang Bai +3
Recent LLMs have demonstrated promising ability in solving finance related problems. However, applying LLMs in real-world finance application remains challenging due to its high ri…
PuzzlePlex: Benchmarking Foundation Models on Reasoning and Planning with Puzzles
Yitao Long, Yuru Jiang, Hongjun Liu +6
This work investigates the reasoning and planning capabilities of foundation models and their scalability in complex, dynamic environments. We introduce PuzzlePlex, a benchmark des…
FinLFQA: Evaluating Attributed Text Generation of LLMs in Financial Long-Form Question Answering
Yitao Long, Tiansheng Hu, Yilun Zhao +2
Large Language Models (LLMs) frequently hallucinate to long-form questions, producing plausible yet factually incorrect answers. A common mitigation strategy is to provide attribut…
SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded Tasks
Yilun Zhao, Kaiyan Zhang, Tiansheng Hu +15
We present SciArena, an open and collaborative platform for evaluating foundation models on scientific literature-grounded tasks. Unlike traditional benchmarks for scientific liter…