6 papers
Think Deep, Not Just Long: Measuring LLM Reasoning Effort via Deep-Thinking Tokens
Wei-Lin Chen, Liqian Peng, Tian Tan +5
Large language models (LLMs) have demonstrated impressive reasoning capabilities by scaling test-time compute via long Chain-of-Thought (CoT). However, recent findings suggest that…
QO-Bench: Diagnosing Query-Operator-Preserving Retrieval over Typed Event Tuples
Mengao Zhang, Xiang Yang, Chang Liu +2
Many real-world questions over business, legal, and scientific corpora are natural-language versions of database-style queries over records latent in text. Existing retrieval-augme…
FinDeepResearch: Evaluating Deep Research Agents in Rigorous Financial Analysis
Fengbin Zhu, Xiang Yao Ng, Ziyang Liu +19
Deep Research (DR) agents, powered by advanced Large Language Models (LLMs), have recently garnered increasing attention for their capability in conducting complex research tasks.…
LikeBench: Evaluating Subjective Likability in LLMs for Personalization
Md Awsafur Rahman, Adam Gabrys, Doug Kang +3
A personalized LLM should remember user facts, apply them correctly, and adapt over time to provide responses that the user prefers. Existing LLM personalization benchmarks are lar…
FAITH: A Framework for Assessing Intrinsic Tabular Hallucinations in Finance
Mengao Zhang, Jiayu Fu, Tanya Warrier +3
Hallucination remains a critical challenge for deploying Large Language Models (LLMs) in finance. Accurate extraction and precise calculation from tabular data are essential for re…
Not All Documents Are What You Need for Extracting Instruction Tuning Data
Chi Zhang, Huaping Zhong, Hongtao Li +11
Instruction tuning improves the performance of large language models (LLMs), but it heavily relies on high-quality training data. Recently, LLMs have been used to synthesize instru…