3 papers
cs.RO2026
Neural Introspection Gating for Adaptive KV-Cache Reuse in Vision-Language-Action Models
Zhijie Wu, Kento Kawaharazuka, Kei Okada
Vision-Language-Action(VLA) models map camera images and language instructions directly to motor commands through a single autoregressive transformer. In real-time control, they st…
cs.CV2026
MERG3R: A Divide-and-Conquer Approach to Large-Scale Neural Visual Geometry
Leo Kaixuan Cheng, Abdus Shaikh, Ruofan Liang +3
Recent advancements in neural visual geometry, including transformer-based models such as VGGT and Pi3, have achieved impressive accuracy on 3D reconstruction tasks. However, their…
cs.CV2024
Representing Animatable Avatar via Factorized Neural Fields
Chunjin Song, Zhijie Wu, Bastian Wandt +2
For reconstructing high-fidelity human 3D models from monocular videos, it is crucial to maintain consistent large-scale body shapes along with finely matched subtle wrinkles. This…