4 papers · 1 filter
PlanningBench: Generating Scalable and Verifiable Planning Data for Evaluating and Training Large Language Models
Ziliang Zhao, Zenan Xu, Shuting Wang +7
Planning is a fundamental capability for large language models (LLMs) because such complex tasks require models to coordinate goals, constraints, resources, and long-term consequen…
AgentFugue: Agent Scaling for Long-Horizon Tasks through Collective Reasoning
Yuyang Hu, Hongjin Qian, Shuting Wang +5
Recent progress on long-horizon agentic tasks has been driven largely by scaling up individual agents through stronger models, better tools, and more effective scaffolding. In cont…
SAM: State-Adaptive Memory for Long-Horizon Reasoning Agent
Yuyang Hu, Hongjin Qian, Shuting Wang +5
Long-horizon agentic reasoning requires large language models to act over long interaction histories containing thoughts, tool calls, observations, and partial conclusions. The cha…
Respond Beyond Language: A Benchmark for Video Generation in Response to Realistic User Intents
Shuting Wang, Yunqi Liu, Zixin Yang +3
Querying generative AI models, e.g., large language models (LLMs), has become a prevalent method for information acquisition. However, existing query-answer datasets primarily focu…