7 papers
Towards Understanding On-Policy Distillation through the Lens of Test-Time Scaling
Xinmu Ge, Zizhuo Zhang, Yu Huang +9
On-policy distillation (OPD) has emerged as a promising post-training technique for enhancing LLM reasoning. It is commonly believed to enable the student model to distill knowledg…
FinReportBench: Measuring and Improving Institution-Grade Financial Report Generation
Yinghao Tang, Tan Zhenwei, Yiyao Wang +4
Large language models can produce fluent financial analysis, but fluency alone does not establish whether a report is suitable for institutional delivery. We introduce FinReportBen…
ReportLogic: Evaluating Logical Quality in Deep Research Reports
Jujia Zhao, Zhaoxin Huan, Zihan Wang +4
Users increasingly rely on Large Language Models (LLMs) for Deep Research, using them to synthesize diverse sources into structured reports that support understanding and action. I…
BOSE: A Systematic Evaluation Method Optimized for Base Models
Hongzhi Luan, Changxin Tian, Zhaoxin Huan +4
This paper poses two critical issues in evaluating base models (without post-training): (1) Unstable evaluation during training: in the early stages of pre-training, the models lac…
Short-Path Prompting in LLMs: Analyzing Reasoning Instability and Solutions for Robust Performance
Zuoli Tang, Junjie Ou, Kaiqin Hu +6
Recent years have witnessed significant progress in large language models' (LLMs) reasoning, which is largely due to the chain-of-thought (CoT) approaches, allowing models to gener…
DAgent: A Relational Database-Driven Data Analysis Report Generation Agent
Wenyi Xu, Yuren Mao, Xiaolu Zhang +4
Relational database-driven data analysis (RDB-DA) report generation, which aims to generate data analysis reports after querying relational databases, has been widely applied in fi…