15 citations · 79 across the 42 of their papers we have counts for
5 papers · 1 filter
TIR-Agent: Training an Explorative and Efficient Agent for Image Restoration
Guoli Jia, Yisheng Zhang, Haote Hu +11
Vision-language agents that orchestrate specialized tools for image restoration (IR) have emerged as a promising method, yet most existing frameworks operate in a training-free man…
AdsQA: Towards Advertisement Video Understanding
Xinwei Long, Kai Tian, Peng Xu +10
Large language models (LLMs) have taken a great step towards AGI. Meanwhile, an increasing number of domain-specific problems such as math and programming boost these general-purpo…
Self-Reflective Reinforcement Learning for Diffusion-based Image Reasoning Generation
Jiadong Pan, Zhiyuan Ma, Kaiyan Zhang +2
Diffusion models have recently demonstrated exceptional performance in image generation task. However, existing image generation methods still significantly suffer from the dilemma…
Retrieval-Augmented Visual Question Answering via Built-in Autoregressive Search Engines
Xinwei Long, Zhiyuan Ma, Ermo Hua +3
Retrieval-augmented generation (RAG) has emerged to address the knowledge-intensive visual question answering (VQA) task. Current methods mainly employ separate retrieval and gener…
Efficient Diffusion Models: A Comprehensive Survey from Principles to Practices
Zhiyuan Ma, Yuzhu Zhang, Guoli Jia +7
As one of the most popular and sought-after generative models in the recent years, diffusion models have sparked the interests of many researchers and steadily shown excellent adva…