12 papers
ProcArena: A Multi-Scenario Benchmark for LLMs on Direct and Interactive PL/SQL Development from Natural Language
Hang Zhang, Chaokun Wang, Yuzhi Pan +12
Large language models (LLMs) have shown strong potential for translating natural-language (NL) requirements into PL/SQL programs, attracting increasing attention from the database…
SLATE: Are AI-Generated Slides Educationally Effective? A Benchmark for Language Teaching Quality and Learner Knowledge Acquisition
Jingzhuo Wu, Jiajun Zhang, Liu Yi +4
LLMs have achieved remarkable capabilities in generating language teaching slides. However, a critical mismatch persists between visual polish and actual instructional effectivenes…
Gradients Know What Outcomes Don't: Unlocking Reinforcement Learning for LLM Reasoning with Gradient-Aligned Rewards
Leqi Zheng, Jinbo Su, Fang Niu +8
Reinforcement learning from verifiable rewards (RLVR) drives chain-of-thought reasoning in large language models, yet its binary outcome reward cannot distinguish among correct tra…
SciLENS: RL-Driven Autonomous Agents for Scientific Localized Evidence Navigation and Synthesis
Leqi Zheng, Jinbo Su, Yuying Li +10
Scientific literature synthesis agents increasingly rely on proprietary online services, limiting reproducibility, privacy, and offline deployment. To address this challenge, we in…
JPO: Juris Policy Optimization for Structured Legal Reasoning in Criminal Judgment Prediction
Zhaolu Kang, Yantao Liu, Tailong Luo +14
Criminal judgment prediction requires models to infer statutory articles, charges, and sentencing outcomes from case facts. Unlike standard classification tasks, it involves a stru…
Filter, Then Reweight: Rethinking Optimization Granularity in On-Policy Distillation
Yuying Li, Leqi Zheng, Yongzi Yu +6
On-Policy distillation (OPD) in large language models is shifting from full-trace KL supervision toward more selective training paradigms. Recent OPD methods increasingly focus on…