2 papers
cs.CL2026
Self-Verification Dilemma: Experience-Driven Suppression of Overused Checking in LLM Reasoning
Quanyu Long, Kai Jie Jiang, Jianda Chen +3
Large Reasoning Models (LRMs) achieve strong performance by generating long reasoning traces with reflection. Through a large-scale empirical analysis, we find that a substantial f…
q-fin.GN2025
Standard Benchmarks Fail -- Auditing LLM Agents in Finance Must Prioritize Risk
Zichen Chen, Jiaao Chen, Jianda Chen +1
Standard benchmarks fixate on how well large language model (LLM) agents perform in finance, yet say little about whether they are safe to deploy. We argue that accuracy metrics an…