10 papers
3D Consistency Optimization for Self-Supervised Monocular Video Depth Estimation
Yuanye Liu, Ke Zhang, Junzhe Jiang +3
Reliable monocular video depth estimation is crucial for downstream 3D reasoning and embodied AI in endoscopic navigation. However, existing self-supervised approaches typically tr…
On-Policy Distillation with Best-of-N Teacher Rollout Selection
Ke Zhang, Yunjie Tian, Dongdi Zhao +4
On-policy distillation (OPD), which supervises a student on its own sampled trajectories, has emerged as a data-efficient post-training method for improving reasoning while avoidin…
Progressive Online Video Understanding with Evidence-Aligned Timing and Transparent Decisions
Kecheng Zhang, Zongxin Yang, Mingfei Han +6
Visual agents operating in the wild must respond to queries precisely when sufficient evidence first appears in a video stream, a critical capability that is overlooked by conventi…
From Imitation to Intuition: Intrinsic Reasoning for Open-Instance Video Classification
Ke Zhang, Xiangchen Zhao, Yunjie Tian +3
Conventional video classification models, acting as effective imitators, excel in scenarios with homogeneous data distributions. However, real-world applications often present an o…
GeoSense: Internalizing Geometric Necessity Perception for Multimodal Reasoning
Ruiheng Liu, Haihong Hao, Mingfei Han +4
Advancing towards artificial superintelligence requires rich and intelligent perceptual capabilities. A critical frontier in this pursuit is overcoming the limited spatial understa…
Endless World: Real-Time 3D-Aware Long Video Generation
Ke Zhang, Yiqun Mei, Jiacong Xu +1
Producing long, coherent video sequences with stable 3D structure remains a major challenge, particularly in streaming scenarios. Motivated by this, we introduce Endless World, a r…