5 papers
Tool-R1: Sample-Efficient Reinforcement Learning for Agentic Tool Use
Yabo Zhang, Yihan Zeng, Qingyun Li +3
Large language models (LLMs) have demonstrated strong capabilities in language understanding and reasoning, yet they remain limited when tackling real-world tasks that require up-t…
AnimateAnywhere: Rouse the Background in Human Image Animation
Xiaoyu Liu, Mingshuai Yao, Yabo Zhang +5
Human image animation aims to generate human videos of given characters and backgrounds that adhere to the desired pose sequence. However, existing methods focus more on human acti…
Beyond Static Scenes: Camera-controllable Background Generation for Human Motion
Mingshuai Yao, Mengting Chen, Qinye Zhou +9
In this paper, we investigate the generation of new video backgrounds given a human foreground video, a camera pose, and a reference scene image. This task presents three key chall…
Personalized Image Generation with Deep Generative Models: A Decade Survey
Yuxiang Wei, Yiheng Zheng, Yabo Zhang +4
Recent advancements in generative models have significantly facilitated the development of personalized content creation. Given a small set of images with user-specific concept, pe…
FramePainter: Endowing Interactive Image Editing with Video Diffusion Priors
Yabo Zhang, Xinpeng Zhou, Yihan Zeng +3
Interactive image editing allows users to modify images through visual interaction operations such as drawing, clicking, and dragging. Existing methods construct such supervision s…