12 papers
BaziQA-Benchmark: Evaluating Symbolic and Temporally Compositional Reasoning in Large Language Models
Jiangxi Chen, Qian Liu
We present BaziQA-Benchmark, a standardized benchmark for evaluating symbolic and temporally compositional reasoning in large language models. The benchmark is derived from 200 pro…
SWE-Dev: Evaluating and Training Autonomous Feature-Driven Software Development
Yaxin Du, Yuzhu Cai, Yifan Zhou +6
Large Language Models (LLMs) have shown strong capability in diverse software engineering tasks. However, feature-driven development, a highly prevalent real-world task that involv…
RC-GRPO: Reward-Conditioned Group Relative Policy Optimization for Multi-Turn Tool Calling Agents
Haitian Zhong, Jixiu Zhai, Lei Song +3
Multi-turn tool calling is challenging for Large Language Models (LLMs) because rewards are sparse and exploration is expensive. A common recipe, SFT followed by GRPO, can stall wh…
Multi-turn Training with Basic Human Feedback Helps Little on LLM Reasoning
Qiang Liu, Wuganjing Song, Zhenzhou Lin +4
The reasoning capabilities of Large Language Models (LLMs) are typically developed through the single-turn reinforcement learning, whereas real-world applications often involve mul…
UPRPRC: Unified Pipeline for Reproducing Parallel Resources -- Corpus from the United Nations
Qiuyang Lu, Fangjian Shen, Zhengkai Tang +4
The quality and accessibility of multilingual datasets are crucial for advancing machine translation. However, previous corpora built from United Nations documents have suffered fr…
PIPA: Preference Alignment as Prior-Informed Statistical Estimation
Junbo Li, Zhangyang Wang, Qiang Liu
Offline preference alignment for language models such as Direct Preference Optimization (DPO) is favored for its effectiveness and simplicity, eliminating the need for costly reinf…