3 papers
cs.CV2026
Iterative Tool Usage Exploration for Multimodal Agents via Step-wise Preference Tuning
Pengxiang Li, Zhi Gao, Bofei Zhang +8
Multimodal agents, which integrate a controller e.g., a vision language model) with external tools, have demonstrated remarkable capabilities in tackling complex multimodal tasks.…
cs.CV2026
MILR: Improving Multimodal Image Generation via Test-Time Latent Reasoning
Yapeng Mi, Yanpeng Zhao, Hengli Li +6
Reasoning-augmented machine learning systems have shown improved performance in various domains, including image generation. However, existing reasoning-based methods for image gen…
cs.CV2025
Building LLM Agents by Incorporating Insights from Computer Systems
Yapeng Mi, Zhi Gao, Xiaojian Ma +1
LLM-driven autonomous agents have emerged as a promising direction in recent years. However, many of these LLM agents are designed empirically or based on intuition, often lacking…