From the 1 of 17 linked papers with an AI index.
17 papers
DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment
Xinyu Geng, Xuanhua He, Sixiang Chen +7
The paper introduces DeepSearch-World, a deterministic, verifiable web environment, and DeepSearch-Evolve, a self‑distillation framework that lets web search agents improve from th…
OPSD-V: On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators
Hongyu Liu, Chun Wang, Feng Gao +6
We propose OPSD-V, an on-policy self-distillation paradigm for post-training few-step autoregressive (AR) video diffusion models. Existing few-step AR video generators can produce…
ICDepth: Taming Video Diffusion Models for Video Depth Estimation via In-Context Conditioning
Xuanhua He, Jiaxin Xie, Mingzhe Zheng +1
Monocular video depth estimation requires temporal consistency, geometric accuracy, and generalization across diverse scenarios, yet existing methods struggle to achieve all three…
LookWise: Knowing When and Where to Look for Fine-Grained Visual Reasoning in Multimodal Large Language Models
Yuxiang Shen, Hailong Huang, Zhenkun Gao +6
Multimodal Large Language Models (MLLMs) are shifting towards "Thinking with Images" by actively exploring image details. While effective, large-scale training is computationally e…
GenEvolve: Self-Evolving Image Generation Agents via Tool-Orchestrated Visual Experience Distillation
Sixiang Chen, Zhaohu Xing, Tian Ye +7
Open-ended image generation is no longer a simple prompt-to-image problem. High-quality generation often requires an agent to combine a model's internal generative ability with ext…
InsEdit: Towards Instruction-based Visual Editing via Data-Efficient Video Diffusion Models Adaptation
Zhefan Rao, Bin Zou, Haoxuan Che +5
Instruction-based video editing is a natural way to control video content with text, but adapting a video generation model into an editor usually appears data-hungry. At the same t…