collaborators

6 papers

cs.LG2026

Decoupling Reasoning and Confidence: Resurrecting Calibration in Reinforcement Learning from Verifiable Rewards

Zhengzhao Ma, Xueru Wen, Boxi Cao +6

Reinforcement Learning from Verifiable Rewards (RLVR) significantly enhances large language models (LLMs) reasoning but severely suffers from calibration degeneration, where models…

cs.SE2026

ScaleBox: Enabling High-Fidelity and Scalable Code Verification for Large Language Models

Jiasheng Zheng, Xin Zheng, Boxi Cao +8

Code sandboxes have emerged as a critical infrastructure for advancing the coding capabilities of large language models, providing verifiable feedback for both RL training and eval…

cs.AI2026

Empowering Small Language Models with Factual Hallucination-Aware Reasoning for Financial Classification

Han Yuan, Yilin Wu, Li Zhang +1

Small language models (SLMs) are increasingly used for financial classification due to their fast inference and local deployability. However, compared with large language models, S…

cs.CL2025

Quantifying the Impact of Structured Output Format on Large Language Models through Causal Inference

Han Yuan, Yue Zhao, Li Zhang +2

Structured output from large language models (LLMs) has enhanced efficiency in processing generated information and is increasingly adopted in industrial applications. Prior studie…

cs.CL2025

Extract, Match, and Score: An Evaluation Paradigm for Long Question-context-answer Triplets in Financial Analysis

Bo Hu, Han Yuan, Vlad Pandelea +3

The rapid advancement of large language models (LLMs) has sparked widespread adoption across diverse applications, making robust evaluation frameworks crucial for assessing their p…

cs.AI2025

Exploring the Reliability of Self-explanation and its Relationship with Classification in Language Model-driven Financial Analysis

Han Yuan, Li Zhang, Zheng Ma

Language models (LMs) have exhibited exceptional versatility in reasoning and in-depth financial analysis through their proprietary information processing capabilities. Previous re…