8 papers
Can Large Language Models Generalize Procedures Across Representations?
Fangru Lin, Valentin Hofmann, Xingchen Wan +4
Large language models (LLMs) are trained and tested extensively on symbolic representations such as code and graphs, yet real-world user tasks are often specified in natural langua…
Visual Planning: Let's Think Only with Images
Yi Xu, Chengzu Li, Han Zhou +4
Recent advancements in Large Language Models (LLMs) and their multimodal extensions (MLLMs) have substantially enhanced machine reasoning across diverse tasks. However, these model…
Agentic Policy Optimization via Instruction-Policy Co-Evolution
Han Zhou, Xingchen Wan, Ivan VuliÄ +1
Reinforcement Learning with Verifiable Rewards (RLVR) has advanced the reasoning capability of large language models (LLMs), enabling autonomous agents that can conduct effective m…
Multi-Agent Design: Optimizing Agents with Better Prompts and Topologies
Han Zhou, Xingchen Wan, Ruoxi Sun +5
Large language models, employed as multiple agents that interact and collaborate with each other, have excelled at solving complex tasks. The agents are programmed with prompts tha…
VISTA: A Test-Time Self-Improving Video Generation Agent
Do Xuan Long, Xingchen Wan, Hootan Nakhost +3
Despite rapid advances in text-to-video synthesis, generated video quality remains critically dependent on precise user prompts. Existing test-time optimization methods, successful…
Maestro: Self-Improving Text-to-Image Generation via Agent Orchestration
Xingchen Wan, Han Zhou, Ruoxi Sun +4
Text-to-image (T2I) models, while offering immense creative potential, are highly reliant on human intervention, posing significant usability challenges that often necessitate manu…