15 papers
Routing Before Looking: Query-Adaptive Evidence Acquisition for Long-form Video Understanding
Tianyue Wang, Xuying Wu, Yuxiang Ma +7
Long-form video understanding remains challenging for video agents due to the mismatch between query demands and evidence acquisition strategies. Although recent planning-before-pe…
DanceOPD: On-Policy Generative Field Distillation
Wei Zhou, Xiongwei Zhu, Zelin Xu +8
Modern image generation demands a single model that unifies diverse capabilities, including text-to-image (T2I), local editing, and global editing. However, these capabilities are…
Optimizing Visual Generative Models via Distribution-wise Rewards
Ruihang Li, Mengde Xu, Shuyang Gu +4
Conventional reinforcement learning strategies for visual generation typically employ sample-wise reward functions, yet this practice frequently results in reward hacking that degr…
OSGNet with MLLM Reranking @ Ego4D Episodic Memory Challenge 2026
Yisen Feng, Leigang Qu, Haoyu Zhang +5
In this report, we present our champion solutions for the Natural Language Queries and GoalStep tracks of the Ego4D Episodic Memory Challenge at CVPR 2026. Both tracks require accu…
AI for Auto-Research: Roadmap & User Guide
Lingdong Kong, Xian Sun, Wei Chow +17
AI-assisted research is crossing a threshold: fully automated systems can now generate research papers for as little as $15, while long-horizon agents can execute experiments, draf…
Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives
Thong Nguyen, Yi Bin, Junbin Xiao +6
Humans use multiple senses to comprehend the environment. Vision and language are two of the most vital senses since they allow us to easily communicate our thoughts and perceive t…