3 papers
cs.SE2026
FLARE: Fine-Grained Diagnostic Feedback for LLM Code Refinement
Yinsheng Yao, Hongxiang Zhang, Weixi Tong +1
Large language models often generate code with bugs. Existing methods rely on feedback signals such as test failures and self-critiques to iteratively refine the generated code. Su…
cs.CL2026
MAVEN: Multi-Agent Verification-Elaboration Network with In-Step Epistemic Auditing
Yinsheng Yao, Jiehao Tang, Zhaozhen Yang +1
While explicit reasoning trajectories enhance model interpretability, existing paradigms often rely on monolithic chains that lack intermediate verification, allowing early errors…
cs.CL2026
FinReasoning: A Hierarchical Benchmark for Reliable Financial Research Reporting
Yiyun Zhu, Yidong Jiang, Ziwen Xu +4
Large language models (LLMs) are increasingly deployed in financial research workflows, where their role is evolving from single-model assistance for human analysts toward autonomo…