8 papers
Where Do Deep-Research Agents Go Wrong? Span-Level Error Localization in Agent Trajectories
Jiaming Wang, Ziteng Feng, Jiangtao Wu +8
Deep-research agents solve tasks through long trajectories of search, tool use, evidence inspection, and answer synthesis. Evaluation based on final answers shows whether an agent…
PEGRL: Improving Machine Translation by Post-Editing Guided Reinforcement Learning
Yunzhi Shen, Hao Zhou, Xin Huang +3
Reinforcement learning (RL) has shown strong promise for LLM-based machine translation, with recent methods such as GRPO demonstrating notable gains; nevertheless, translation-orie…
DR-Eval: Towards Realistic and Reproducible Deep Research Evaluation
Qianqian Xie, Qingheng Xiong, He Zhu +16
Deep Research Agents (DRAs) aim to solve complex, long-horizon research tasks involving planning, retrieval, multimodal understanding, and report generation, yet their evaluation r…
R3S: Refining and Recovering Reinforcement Signals for Multilingual Understanding and Reasoning
Junxiao Liu, Zhijun Wang, Yixiao Li +6
Large reasoning models often default to English reasoning when processing non-English questions, yet their performance drops substantially when reasoning in the question language.…
Align to the Pivot: Dual Alignment with Self-Feedback for Multilingual Math Reasoning
Chunxu Zhao, Xin Huang, Xue Han +3
Despite the impressive reasoning abilities demonstrated by large language models (LLMs), empirical evidence indicates that they are not language agnostic as expected, leading to pe…
Understanding LLMs' Cross-Lingual Context Retrieval: How Good It Is And Where It Comes From
Changjiang Gao, Hankun Lin, Xin Huang +5
Cross-lingual context retrieval (extracting contextual information in one language based on requests in another) is a fundamental aspect of cross-lingual alignment, but the perform…