1 citations · 1 across the 12 of their papers we have counts for
8 papers · 1 filter
Self-Supervised Visual On-Policy Distillation
Yijiang Li, Yijun Liang, Yunjie Tian +6
Visual on-policy distillation relies heavily on an informative teacher-student asymmetry, through either a larger, stronger teacher or privileged supervision, such as reference ans…
3D Consistency Optimization for Self-Supervised Monocular Video Depth Estimation
Yuanye Liu, Ke Zhang, Junzhe Jiang +3
Reliable monocular video depth estimation is crucial for downstream 3D reasoning and embodied AI in endoscopic navigation. However, existing self-supervised approaches typically tr…
Progressive Online Video Understanding with Evidence-Aligned Timing and Transparent Decisions
Kecheng Zhang, Zongxin Yang, Mingfei Han +6
Visual agents operating in the wild must respond to queries precisely when sufficient evidence first appears in a video stream, a critical capability that is overlooked by conventi…
GeoSense: Internalizing Geometric Necessity Perception for Multimodal Reasoning
Ruiheng Liu, Haihong Hao, Mingfei Han +4
Advancing towards artificial superintelligence requires rich and intelligent perceptual capabilities. A critical frontier in this pursuit is overcoming the limited spatial understa…
Endless World: Real-Time 3D-Aware Long Video Generation
Ke Zhang, Yiqun Mei, Jiacong Xu +1
Producing long, coherent video sequences with stable 3D structure remains a major challenge, particularly in streaming scenarios. Motivated by this, we introduce Endless World, a r…
FreeViS: Training-free Video Stylization with Inconsistent References
Jiacong Xu, Yiqun Mei, Ke Zhang +1
Video stylization plays a key role in content creation, but it remains a challenging problem. Naïvely applying image stylization frame-by-frame hurts temporal consistency and reduc…