6 papers
DrivingDepth: Sparse-Prompted Pixel-wise Scale Correction for Driving Depth Estimation
Chi Huang, Wenhao Zhang, Hang Yin +5
Dense depth estimation for autonomous driving faces a geometry-scale conflict: depth foundation models deliver pixel-aligned dense visual geometry without reliable metric scale, wh…
VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding
Kuanwei Lin, Wenhao Zhang, Ge Li
Video large multimodal models increasingly face a scalability bottleneck: long videos produce excessively long visual-token sequences, which sharply increase memory and latency dur…
Dynamic Pondering Sparsity-aware Mixture-of-Experts Transformer for Event Stream based Visual Object Tracking
Shiao Wang, Xiao Wang, Duoqing Yang +5
Despite significant progress, RGB-based trackers remain vulnerable to challenging imaging conditions, such as low illumination and fast motion. Event cameras offer a promising alte…
FaithFusion: Harmonizing Reconstruction and Generation via Pixel-wise Information Gain
YuAn Wang, Xiaofan Li, Chi Huang +5
In controllable driving-scene reconstruction and 3D scene generation, maintaining geometric fidelity while synthesizing visually plausible appearance under large viewpoint shifts i…
WIPES: Wavelet-based Visual Primitives
Wenhao Zhang, Hao Zhu, Delong Wu +4
Pursuing a continuous visual representation that offers flexible frequency modulation and fast rendering speed has recently garnered increasing attention in the fields of 3D vision…
MTGA: Multi-View Temporal Granularity Aligned Aggregation for Event-Based Lip-Reading
Wenhao Zhang, Jun Wang, Yong Luo +4
Lip-reading is to utilize the visual information of the speaker's lip movements to recognize words and sentences. Existing event-based lip-reading solutions integrate different fra…