1 citations · 2 across the 10 of their papers we have counts for
8 papers · 1 filter
AlgoWorlds: Benchmarking Tool Use for Global Optimization in Algorithmic Worlds
Zixiang Xu, Jiaan Wang, Fandong Meng
Tool-use benchmarks generally evaluate whether an agent completes a workflow using appropriate tools and valid arguments. However, feasibility alone is insufficient in real-world d…
EvoBrowseComp: Benchmarking Search Agents on Evolving Knowledge
Yunhan Wang, Jiaan Wang, Lianzhe Huang +2
Search Agents -- large language models augmented with search tools -- have intensified the need for future-proof evaluation benchmarks. Existing benchmarks such as BrowseComp rely…
SlangDIT: Benchmarking LLMs in Interpretative Slang Translation
Yunlong Liang, Fandong Meng, Jiaan Wang +1
The challenge of slang translation lies in capturing context-dependent semantic extensions, as slang terms often convey meanings beyond their literal interpretation. While slang de…
ExTrans: Multilingual Deep Reasoning Translation via Exemplar-Enhanced Reinforcement Learning
Jiaan Wang, Fandong Meng, Jie Zhou
In recent years, the emergence of large reasoning models (LRMs), such as OpenAI-o1 and DeepSeek-R1, has shown impressive capabilities in complex problems, e.g., mathematics and cod…
An Empirical Study of Many-to-Many Summarization with Large Language Models
Jiaan Wang, Fandong Meng, Zengkui Sun +5
Many-to-many summarization (M2MS) aims to process documents in any language and generate the corresponding summaries also in any language. Recently, large language models (LLMs) ha…
DeepTrans: Deep Reasoning Translation via Reinforcement Learning
Jiaan Wang, Fandong Meng, Jie Zhou
Recently, deep reasoning LLMs (e.g., OpenAI o1 and DeepSeek-R1) have shown promising performance in various downstream tasks. Free translation is an important and interesting task…