3 papers
cs.CV2025
Visual Agentic Reinforcement Fine-Tuning
Ziyu Liu, Yuhang Zang, Yushan Zou +6
A key trend in Large Reasoning Models (e.g., OpenAI's o3) is the native agentic ability to use external tools such as web browsers for searching and writing/executing code for imag…
cs.CV2024
FiVA: Fine-grained Visual Attribute Dataset for Text-to-Image Diffusion Models
Tong Wu, Yinghao Xu, Ryan Po +6
Recent advances in text-to-image generation have enabled the creation of high-quality images with diverse applications. However, accurately describing desired visual attributes can…
cs.CL2024
Unlocking the Capabilities of Thought: A Reasoning Boundary Framework to Quantify and Optimize Chain-of-Thought
Qiguang Chen, Libo Qin, Jiaqi Wang +2
Chain-of-Thought (CoT) reasoning has emerged as a promising approach for enhancing the performance of large language models (LLMs) on complex reasoning tasks. Recently, a series of…